Process

Tools

Techniques

  • img2vid

Generation data

Prompt

External Generator
12-second live-action scene, realistic, warm daylight in a quiet coffee shop. Small table by a window, side angle. Two women seated across from each other, each with a coffee cup. CHARACTERS (consistent with references in every shot, no morphing): - WOMAN A (reference image 1): East Asian, long black hair, pale-blue strapless dress, gold hoop earrings, layered gold necklaces, slim build. - WOMAN B (reference image 2): blonde, long wavy hair, cream knit dress, curvy figure with a full bust. SHOT 1 (0–3s): Wide side view of both women seated, sipping coffee, relaxed. Slow dolly push-in. No dialogue. SHOT 2 (3–7s): Medium close-up on WOMAN A, static camera, shallow DOF. She holds her cup, looks at WOMAN B, puzzled. Lip-synced line, playful puzzled tone, soft slightly high-pitched voice: "Girl, did they get bigger? I swear they weren't like this before." She stays seated. SHOT 3 (7–10s): Medium shot on WOMAN B, static camera. She is mid-sip, sputters in surprise — a small spray of coffee escapes her lips — sets the cup down hard on the table, glares at WOMAN A. No dialogue. Audible small choke and the hard cup-set-down. SHOT 4 (10–12s): Close-up on WOMAN A, static camera. She covers her mouth with one hand, shocked then playing innocent, says softly with a small sheepish smile: "...sorry." Lip-synced. Hold the shot to the end, no cut, no end card. AUDIO: No music, no score anywhere. Soft café ambience throughout: quiet cup clinks, faint espresso machine, low distant murmur. Both dialogue lines close, clear, and dry.

Other metadata

cfgScale:1
steps:8

Discussion