We Made Two Short Films With Dialogue Using One Video Model
Two complete short films — a magic-noir tragedy and a stranded-crew sci-fi story — made with 16 structured H3 prompts: one verbatim IDENTITY LOCK per shot for character consistency, dialogue written into the prompt for generation-time lip-sync, and music mixed in afterwards. The workflow, prompts, and render numbers.
We made two complete short films with dialogue — Ember, a magic-noir tragedy set in a desert city, and Hollow Orbit, the story of a crashed survey ship on a bioluminescent tidally-locked planet — using nothing but our H3 video model on our own GPU stack. Each film is 8 shots × 5 seconds: 16 shots total, and every shot is a single structured prompt. H3 has no memory between shots, so character consistency lives in the prompt itself, and dialogue is generated by the model, with lip-sync happening in generation. This article is the whole workflow: the identity technique, the dialogue format, why music never enters the model, and what the renders cost.
Two films, sixteen prompts
Both films were planned as shot lists first and prompts second — one prompt per shot, rendered sequentially on one local GPU worker. Ember is a magic-noir tragedy: a street sorceress steals a forbidden ember to save her dying brother, the guard who loves her takes a killing bolt for her, and she burns her own spark to save him, scattering into embers at sunrise. Strife, love, loss — forty seconds of it. Hollow Orbit strands a survey ship on a tidally-locked planet: the alien Choir rebuilds their engine but demands a resident mind, co-pilot Lena chooses to stay and becomes the helm, and the valley blooms in her tattoo geometry as the ship escapes.
PROMPT
[SHOT e1 — verbatim from the spec, opening truncated]
PHOTOREAL CINEMATIC FANTASY, anamorphic 2.39:1, warm ember palette against
cold indigo pre-dawn. IDENTITY LOCK: SERA, woman mid-20s, sun-bronzed olive
skin, black braided crown hair with loose windblown strands, amber eyes,
glowing copper sigil tattoos spiraling up both forearms, crimson open-backed
wrap top baring her shoulder blades, long charcoal skirt with a high slit...
SCENE: She stands on the ribbed clay rooftop of a teetering desert city at
first light, dealing a fan of tarot-sized cards whose faces smolder like
coals... AUDIO: dry wind over clay, distant bell-goat chimes, the soft hiss
of the burning card. No dialogue, no music, no subtitles, no logos.OUTPUT — GENERATED WITH MANIFOLDGEN
Ember — 8 shots × 5 seconds, generated and cut end to end
Shot 1 establishes geography and the theft in one prompt. The IDENTITY LOCK pins Sera while everything underneath it — location, action, camera, light, audio — is new per shot; the remaining seven prompts change only those variables.
PROMPT
[SHOT h1 — verbatim from the spec, opening truncated]
PHOTOREAL SCIENCE FICTION, anamorphic 2.39:1, awe-heavy establishing.
SCENE: the survey ship MERIDIAN LARK, a hundred-meter wedge, crashed nose-down
in a valley of glassy violet coral under a locked twilight sky; a banded gas
giant hangs fixed on the horizon; aurora curtains ripple green-violet...
ACTION 0-5s: two tiny suited figures descend the rope from the airlock onto
coral that glows faintly brighter around each of their footfalls, light
rippling outward like pond rings... No dialogue, no music, no subtitles.OUTPUT — GENERATED WITH MANIFOLDGEN
Hollow Orbit — 8 shots × 5 seconds, dialogue generated by the model
The whole film is eight such prompts. The ship, the valley, and the Choir reappear across shots because their descriptors repeat verbatim — the same discipline the IDENTITY LOCK applies to faces.
How character consistency works
Each H3 generation is independent: shot 7 has no knowledge of shot 6. Consistency is therefore a prompt discipline, not a model feature. Every shot prompt opens with the same IDENTITY LOCK block, pasted verbatim, that pins everything about a character that must never change. Here is the truncated SERA lock from Ember's spec:
IDENTITY LOCK, repeat every shot:
SERA, woman mid-20s, sun-bronzed olive skin, black braided crown hair with
loose windblown strands, amber eyes, glowing copper sigil tattoos spiraling
up both forearms, crimson open-backed wrap top baring her shoulder blades,
long charcoal skirt with a high slit, brass bangles, wrapped barefoot sandals.
KADE, city guard about 30, brown skin, close-cropped black hair, trimmed
stubble, a scar over his right brow, dented bronze scale armor over an
indigo tunic, crimson sash at the waist.
[... repeated verbatim at the top of all eight prompts ...]The lock fixes face, hair, wardrobe, jewelry, even how her magic manifests — and it is pasted into all eight of Ember's prompts without changing a word. The shot prompt underneath changes only what should change: location, action, camera, light, sound. Drift starts the moment you reword the lock; a "black braided crown" that becomes "loose beach waves" in shot 6 is a different character.
Dialogue that actually lip-syncs
Lines are written into the shot prompt with their exact wording, and H3 performs them — lip-sync happens during generation, not in a post pass. Two rules hold every dialogue shot together: name the speaker, and mark the words as fixed. Ember's last line to the guard, as she burns her spark at dawn:
[SHOT e8 — sacrifice finale, sunrise]
ACTION AND DIALOGUE: 0-2.5s she presses the guttering EMBER into his chest;
its light floods his wound; her own sigils stream out like molten thread into
him. SERA, a whisper, smiling: "Live brighter." 2.5-4s: his eyes open, gasp;
hers go softly dark as her body loosens into drifting embers and ash-light.And Hollow Orbit's decision beat, once the Choir has rebuilt the engine and named its price:
[SHOT h6 — cockpit conflict]
ACTION AND DIALOGUE: 0-1.5s IVAR slams the console. IVAR, raw: "We fly OUT."
1.5-3.5s LENA, steady, terrified, lip-synced: "Only a mind can hold the helm
it's offering. A machine can't carry it." 3.5-5s: hold on the two-shot, storm
light flaring between them, neither blinking.Both lines are short on purpose: a close-up, one clause, the exact words in quotes, and an Audio: line carrying voice and diegetic sound only. The "no score" instruction is not decoration — it is what makes the next section possible.
Music stays out of the model
H3 renders dialogue, ambience, and effects in generation, which means a score baked into the same pass would fight the dialogue track inside the model's own audio. We keep music out entirely and treat it as a mix problem.
Instrumental beds are generated separately with our music model and mixed under the dialogue in ffmpeg, sidechain-ducked so the bed dips whenever a character speaks. Every Audio: line in a shot prompt lists diegetic sound only — voices, wind, hull hum. Music never enters generation.
What it costs to run
| Stage | How it ran |
|---|---|
| Shot rendering | 16 shots — 8 per film, 5 seconds each, one structured prompt per shot |
| Hardware | Our own GPU stack: one local GPU worker, shots rendered sequentially |
| Time per shot | Roughly 4–10 minutes per 5-second shot |
| Character consistency | IDENTITY LOCK block repeated verbatim in every prompt |
| Dialogue | Written into the prompt with exact wording, generated and lip-synced by H3 |
| Music | Instrumental beds from our music model, mixed under dialogue in ffmpeg |
The same stack is open in the studio: direct shots at /studio, score them with the generator at /tools/music-generator, and keep every asset searchable on this site. Write one IDENTITY LOCK, paste it into every prompt verbatim, put the exact spoken words in quotes, and keep music out of the model — that is the entire method behind both films.
Try these prompts on your own shot.
Every example on this page was generated with ManifoldGen. Open the Studio and run the same structure on your idea.