Blog

How To Make a Game Trailer with AI (Full Workflow + Prompts)

HiggsfieldAug 10, 202613 min
How To Make a Game Trailer with AI (Full Workflow + Prompts)

To make a game trailer with AI, you need three things: locked character and location references, a shot-by-shot script, and a video model that can animate those references consistently. This guide walks through the full workflow, from character sheets to final cut, using Cinema Studio for generation and Marketing Studio to turn the same assets into ad creatives afterward.

Key takeaways

  • A trailer this size runs on surprisingly few locked references: this example uses 3 designed characters and 7 locations, reused across every shot instead of regenerated per scene.

  • Not every look needs its own asset. A character in a different outfit, or a location's own built-in equipment, can be described directly inside a scene's prompt when that look only needs to hold within one storyline thread.

  • Plan the full shot list before generating anything. This trailer breaks into 4 acts across two 30-second generations, about a minute total, decided on paper first.

  • Once the trailer is done, the same locked files go straight into Marketing Studio for a batch of ad variants, no rebuilding required.

What is a game trailer, and what type does this guide build?

A game trailer is a short promotional video built to sell a game. There are a few recurring types: teaser and announcement trailers that introduce a game without revealing much, cinematic or story trailers that dramatize character and world the way a short film would, gameplay trailers that show actual mechanics and moment-to-moment feel, and launch trailers timed to release to convert interest into wishlists and purchases.

This guide builds a cinematic trailer: a fictional example built for a video game named LETA that leans entirely on character, world and atmosphere rather than gameplay footage, since that's the format an asset-first AI workflow is best suited for. If your own game needs a gameplay trailer instead, the same locked-reference workflow still applies, only the shot list changes from dramatized scenes to actual mechanics.

What do you need before generating a game trailer with AI?

Before generating anything, you need:

  • A concept and a full shot-by-shot script. A list of scenes, dialogue, sound cues and cuts included.

  • A written description for each character, detailed enough that someone else could picture them from the text alone. That's what keeps a face consistent later.

  • A locked reference sheet per character, generated once so the same face holds across every shot.

  • Variants for any character who changes look, like a patient in a hospital gown instead of everyday clothes, generated from the base character rather than written as a new character from scratch.

Once all of that exists, generate the script scene by scene, in order.

How do you storyboard a game trailer before generating anything?

Write the full trailer as a shot list before generating a single frame, cuts, dialogue and sound included, then generate against that list instead of improvising order later. Most trailers, whatever the story, break into a similar four-part arc:

  • The hook: something unresolved in the first few seconds, before the trailer explains anything.

  • The rise: world and characters established, tone still light or neutral.

  • The turn: tone shifts, tension escalates, stakes become clear.

  • The climax: ends unresolved rather than tidy, then the title card.

How do you make a game trailer on Higgsfield?

Inside Higgsfield, each part of the workflow above maps to a specific tool, all inside Cinema Studio. Characters and locations are generated in Soul Cinema, so a face or a room can be locked once and reused across every shot. Every video shot, from a single close-up to a full multi-shot scene, is generated in Seedance 2.5. Cinema Studio is how Higgsfield handles all of it end to end, instead of switching between separate tools for each part of the trailer.

Cinema Studio also gives you deterministic camera control, and for a game trailer that matters more than almost anything else. You can pick a virtual camera body, lens and focal length, choose specific movements like a dolly, a crane rise or a crash zoom, and stack several camera moves inside one shot, instead of hoping the model improvises something usable. Every lens lock and push-in in the trailer prompts below is written against that control, which is why the shot descriptions can be as specific as they are.

Everything below, the characters, the locations and the trailer itself, was built this way.

What's the story behind LETA?

LETA is a fictional near-future detective game built around a memory-editing procedure that's licensed for therapy and quietly used for cover-ups. Kai Weil runs the procedure, clinical and unreadable by training. His current client is Leta Moren, an heiress whose husband Damen Voss ordered the extraction to erase her memory of his affair and buy time before she finds out he's after her inheritance. Partway into the extraction, Kai finds his own face inside a memory he shouldn't recognize, and the deeper he goes, the more it looks like Leta's family history hides something Damen needs buried for good.

Mapped onto the four-part arc above, LETA's trailer runs: a dark, unresolved teaser, a fast rise through Leta's warm years, a turn where the relationship and the family history both start cracking, and a climax in the clinic that ends on a standoff, not an answer.

Who are the characters in LETA?

Three designed characters carry the trailer, all generated in Soul Cinema. Because the story moves across two decades, each of them is also generated at a younger age, and Leta's parents get one shared reference, all built on the same base descriptions below. What never needs its own asset is a change of clothes: Leta wears a different outfit in every scene, and each one is written into the scene prompt on top of the same locked face.

Kai Weil, the memory editor (@kai)

What it does: generates Kai's reference sheet, the version of his face used in every clinic and memory shot.

Character sheet of one man, three panels side by side, plain grey studio backdrop, thin dark dividers. Left: full-body front, head to toe, standing straight. Center: full-body back, same pose. Right: close-up of face and shoulders, calm unreadable expression. Kai Weil, 30: fair skin with a cool under-slept pallor, dark brown almost black hair cut short and neatly, dark brown eyes, lean and narrow through the shoulders with visible tendon at the wrists and neck. His face carries almost no mimic lines, no crow's feet, no forehead lines, no nasolabial folds, a fully neutral trained mask, strangely smooth for his age. A thin pale horizontal scar, 15mm, at the left temple just inside the hairline. Outfit: a clean white clinical coat worn open over a dark charcoal-grey crew-neck base layer and slim dark charcoal trousers, no logos, a thin matte dark grey interface cuff on the right wrist. Lighting: neutral soft even wraparound, near-shadowless, 5600K, no color cast, readable skin texture, eye catchlights. LOCKS: one identical man in all three views, same identity, outfit, hair, proportions; the temple scar visible in the close-up panel; identical backdrop and light.

Damen Voss, the husband (@damen)

What it does: generates Damen's adult reference sheet, used in every present-day beat, with his suit and frown lines locked so they never drift between cuts.

Character sheet of one man, three panels side by side, plain grey studio backdrop, thin dark dividers. Left: full-body front, head to toe, standing straight. Center: full-body back, same pose. Right: close-up of face and shoulders, cold controlled expression. Damen Voss, 32: tall and solidly built through the chest with the beginnings of softness at the waist, ash-blond hair with clear grey at both temples combed back with visible product, fair skin with a faint flush and broken capillaries across the nose and cheeks, pale grey-blue cold eyes beneath a heavy brow. His face is carved by frowning, two deep permanent horizontal lines across the forehead and a third shorter one between the brows, present even at rest, plus pronounced nasolabial folds. Outfit: an expensive black wool suit, jacket open and unbuttoned, over a white dress shirt with the top two buttons undone and no tie, a heavy steel watch and a plain gold wedding band. Lighting: neutral soft even wraparound, near-shadowless, 5600K, no color cast, readable skin texture. LOCKS: one identical man in all three views, same identity, outfit, hair, proportions; identical backdrop and light.

Leta Moren, the heiress (@leta)

What it does: generates Leta's base reference sheet in her everyday clothes, used in every memory scene outside the clinic.

Character sheet of one woman, three panels side by side, plain grey studio backdrop, thin dark dividers. Left: full-body front, head to toe, standing straight. Center: full-body back, same pose. Right: close-up of face and shoulders, soft melancholic expression. Leta Moren, 29: medium build, light chestnut hair falling loose past the shoulders in a soft wave with loose face-framing strands, golden-amber eyes with visible gold flecks, faint freckling across the nose bridge, a small dark mole below the outer corner of the left eye, soft natural makeup, small gold stud earrings, a slim gold band on her finger. Outfit: plain neutral everyday clothing, no logos, no pattern. Lighting: neutral soft even wraparound, near-shadowless, 5600K, no color cast, readable skin texture, eye catchlights. LOCKS: one identical woman in all three views, same identity, hair, proportions, freckling and mole in the same positions; identical backdrop and light.

What locations does the LETA trailer need?

Seven locations carry the story, all generated the same way. A few smaller settings appear in single beats only, the bedroom, the family dining table and the doorway, and those were written straight into their scene prompts instead of being generated as separate assets.

The clinic (@clinic), laid out the same way in every shot so nothing drifts between the wide plan and the close inserts.

Photoreal, 8K, cinematic, bright cold clinical light, almost dental-office in feel. LAYOUT: left side of the room holds the extraction machine console. Center, against the back wall, a reclined chair for the patient. Right side, against the wall, a small waiting area with potted plants and two armchairs. Wide shot, all three zones visible in frame, hard white-blue light, polished pale floor. Empty of people. No text, no real-world brands.

The playground (@playground), generates for the earliest memories.

Photoreal, 8K, cinematic. A quiet neighborhood playground in late afternoon, shot from a 3/4 angle. A wooden sandbox, an empty swing set, warm low golden light, soft nostalgic film grain. Empty of people. No text, no real-world brands.

The beach (@beach), generates for the memory where Leta and her friends run toward the water at sunset.

Photoreal, 8K, cinematic. A wide beach at sunset, shot from the sand looking out, the waterline entering the frame from the middle of the left side rather than the right. The water reflects the low sun and reads gold rather than blue. Warm golden haze, no other landmarks. Empty of people. No text, no real-world brands.

The bar (@bar), generates for the warm night out with friends during Leta's younger years.

Photoreal, 8K, cinematic. A small dim bar interior at night, shot from a 3/4 angle. Warm amber light, a worn wooden counter, soft out-of-focus string lights in the background. Empty of people. No text, no real-world brands.

The school (@school), generates for the classroom where Leta and Damen first meet, used in both halves of the trailer.

Photoreal, 8K, cinematic. A high school classroom, shot toward the back row of single desks, bookshelves lining the left wall, tall windows on the far wall throwing strong rim backlight with visible volumetric light shafts. Warm classroom daylight, backlit haze, gentle bloom along the window edge. Empty of people. No text, no real-world brands.

The parents' office (@parents-office), generates for scenes where Leta breaks down.

Photoreal, 8K, cinematic. A parents' formal home study, shot from a 3/4 angle, dark wood furniture, heavy curtains, cold thin daylight, harder contrast with deep shadow gathering in the corners of the room, an oppressive stillness. Empty of people. No text, no real-world brands.

The backyard (@backyard), the memory the whole trailer is built around.

Photoreal, 8K, cinematic. A family backyard at dusk, shot from a 3/4 angle, tall trees along the fence line with clear gaps between the trunks that a view can pass through, a shovel resting beside a mound of freshly turned earth near the treeline. Fading dusk light shifting from low warm to blue, soft volumetric haze between the trunks. Empty of people. No text, no real-world brands.

How do you keep the same face and location across every generation?

Once every asset above exists, keep it in one place so the same file goes into every shot instead of a fresh face per scene. The handles above are your own naming, useful for finding assets fast. The prompts themselves read references by upload position: the first image you attach becomes @image_1, the second @image_2, and so on. That makes upload order part of the prompt, and the two halves of this trailer use different orders, so each one lists its own sequence. Locations work differently again. In this trailer they were not attached as references at all, they are written out in words inside each scene prompt, with the locked location images used as the visual target while writing those descriptions.

One rule matters more than any other with Seedance: it doesn't process negative prompts. Writing “she doesn't smile” tends to produce the opposite. Describe the state you actually want instead.

How do you keep the same face and location across every generation?

Type

Quantity in this trailer

Model

Credits

≈ USD (at ~20 credits/$1)

Characters

3

Soul Cinema

0.375

~$0.02

Locations

7

Soul Cinema

0.875

~$0.04

Video scenes

2 (30s each)

Seedance 2.5, 720p

~390

~$19.50

Total

~391.25

~$19.56

At roughly 20 credits per dollar, a trailer this size costs approximately $19.56 in credits, and almost all of it goes to video. Still images are effectively free by comparison: all ten character and location references together cost about 1.25 credits, well under one percent of the total. Video pricing scales with length, 195 credits per 30-second generation at 720p, so your own cost depends almost entirely on how much finished runtime you generate. The $1 ≈ 20 credits rate is a rough approximation, not tied to any specific subscription tier, check the pricing page for the exact rate on your plan.

The trailer cut

Two generations carry the full four-act trailer, about a minute end to end.

Both prompts ask for a deliberately rendered look rather than a photographic one: near-photoreal fidelity with the tells of a real-time engine, subsurface sheen on skin, strand-clumped hair, computed bokeh, an even grain. That is a choice. A cinematic game trailer has to read as belonging to a game, and an image that looks photographed reads as a film instead. The character and location references stay photographic, since they only carry identity and layout, and the video prompt sets the final look.

Part 1: the setup (0:00 to 0:30)

What it does: generates the entire first half in one pass, the teaser flashes, the clinic reveal, the interface prompt, Kai's glance, the warm montage and the exchange with Damen. Because references attach by upload order, load them in exactly this sequence: 1 Kai, 2 Leta in her mid-twenties, 3 Leta at 29, 4 Damen, 5 Leta at six, 6 Damen at sixteen.

SCENE CONTEXT Opening half of a psychological sci-fi trailer, rendered as a HIGH-END REAL-TIME GAME CINEMATIC. A memory-extraction operator is about to delete a woman's memories while her husband waits and watches. Fragments of violence flash, the clinic is revealed, the "delete memory" choice appears, and a montage of warm years plays, raising the question of why anyone would erase this. Ends on quiet tension between the operator and the husband. RENDERING STYLE — MANDATORY, APPLIES TO EVERY BEAT: this must read as a top-tier real-time rendered game cinematic in the register of a modern narrative console title — near-photoreal fidelity but unmistakably a rendered engine image rather than photographed footage. Every tell below must be present in every shot. SKIN: full micro-detail — pores, fine lines, moles and stubble all rendered — but with the characteristic engine subsurface look: smooth even translucency beneath the surface and a LIGHT UNIFORM SPECULAR SHEEN across the forehead, the bridge of the nose, the cheekbones and the chin, as though very faintly damp. Slightly waxy, slightly too clean, with no oil variation and no blotchiness. Detail evenly distributed across the whole face with no falloff. EYES: glassy and high-specular, with a clean bright sclera, a crisp perfectly circular catchlight in each eye, a strong wet highlight on the cornea and a sharply defined limbal ring. Slightly too perfect, slightly synthetic. Individually rendered eyelashes. HAIR: rendered in defined strand clumps and card-like groupings rather than fully individual chaotic strands, with a uniform anisotropic sheen running along the length and clean silhouette edges. MATERIALS: physically based shading throughout, every surface with correct roughness and reflectance, visible screen-space reflections on glass, metal, polished floors and screens, and soft ambient occlusion darkening every crease, seam and contact point. All fabric reads as simulated cloth — clean folds, slightly stiff drape, uniform wrinkle logic. IMAGE AND POST: everything critically sharp across the full frame with no lens softness, no field curvature and no corner falloff. Detail and contrast do NOT diminish with distance — backgrounds are as cleanly resolved as foregrounds. Where depth of field is used it is computed and clean with perfectly circular uniform bokeh. Motion blur is present and correct but clean and per-object rather than optical smear. Gentle uniform bloom on the brightest highlights. Subtle chromatic aberration applied evenly as a post effect. A fine even film grain overlay laid uniformly across the whole image. Slightly elevated contrast and clean highlight rolloff. All camera movement is smooth interpolated virtual-camera motion with no handheld shake and no operator instability. FORBIDDEN: no illustration, no anime, no cel shading, no stylised or exaggerated proportions, no painterly look, no concept art, no cartoon, no low-poly, no visible polygon edges, no texture seams, no flat shading, no clay render, no grey-boxed or unfinished surfaces, no flat repeated crowd instances. Fidelity must be maximal and the image must look expensive — only the LOOK is rendered rather than photographic. ASSET REGISTRY (reference handles = upload order; identity is 100% from each reference) @image_1 — KAI WEIL, 30. Operator. Fair skin with a cool under-slept pallor, dark brown almost black hair cut short and neatly, dark brown eyes, lean and narrow through the shoulders with visible tendon at the wrists and neck. HIS FACE CARRIES ALMOST NO MIMIC LINES — no crow's feet, no forehead lines, no nasolabial folds, no frown lines — a fully neutral trained mask, strangely smooth for his age. A thin pale horizontal scar, 15mm, at the left temple just inside the hairline. Wears a clean WHITE clinical coat, open, over a dark charcoal-grey crew-neck base layer and slim dark charcoal trousers. A thin matte dark grey interface cuff on the right wrist. Calm, controlled, unreadable. 100% matches the reference. @image_2 — LETA MOREN, youthful era. 24 to 25. LIGHT CHESTNUT hair pulled back with loose face-framing strands, GOLDEN-AMBER eyes with visible gold flecks, faint freckling across the nose bridge, glam makeup, gold hoops and a thin gold necklace. Luminous and confident. Same face and same eye colour as @image_3. Used only for the youthful social memories. 100% matches the reference. @image_3 — LETA MOREN, present. 29. The same LIGHT CHESTNUT hair, now falling loose past the shoulders in a soft wave, the same GOLDEN-AMBER eyes, the same freckling, soft natural makeup, small gold studs, a slim gold band on her finger. Softer, melancholic. In the clinic she wears a plain pale grey-blue hospital gown, not the clothing from the reference. 100% matches the reference for face and identity. @image_4 — DAMEN VOSS, 32. The husband and antagonist. Tall and solidly built through the chest with the beginnings of softness at the waist. ASH-BLOND HAIR WITH CLEAR GREY AT BOTH TEMPLES, combed back. Fair skin with a faint flush and broken capillaries across the nose and cheeks. Pale grey-blue cold eyes beneath a heavy brow. HIS FACE IS CARVED BY FROWNING — two deep permanent horizontal lines across the forehead and a third shorter one between the brows, present even at rest, plus pronounced nasolabial folds. Wears an EXPENSIVE BLACK WOOL SUIT, jacket open and unbuttoned, over a white dress shirt with the top two buttons undone and NO TIE. A heavy steel watch, a plain gold wedding band. Cold and controlled, with nervous slightly mechanical smiles. 100% matches the reference. @image_5 — LETA as a young child, about 6. Built from @image_3's features: light chestnut hair just past the shoulders, golden-amber eyes, faint freckling across the nose bridge, a small mole below the outer corner of the left eye in the same position. Playground only. @image_6 — DAMEN as a teenager, about 16. Built from @image_4's features but softer and boyish: ash-blond hair, pale grey-blue eyes, no grey and no frown lines yet, in a plain dark school jumper over a pale shirt. School desk only. GLOBAL LOCKS (hold across every beat) Identity is 100% from the referenced images; never redesign a face between cuts. Clinic beats always use the same cold clinical high-key grade so they read as one place. Memory beats use a warm nostalgic grade. Only the two scripted English lines are spoken; all other lips stay still; no narration, no offscreen extra voices. A slow steady heart-monitor beep runs under the clinic beats. NO TEXT anywhere except the one console UI specified below — no logos, no signage, no name badges, no ID cards, no lanyards, no wall labels, no equipment markings. SHOT SEQUENCE — controlled multi-shot, real-time, hard cuts only unless stated. 0.0s–1.5s — TEASER FRAGMENTS Three sub-second flashes, each its own micro-cut: a mouth open mid-scream with no clear words; a plate shattering; a human silhouette in sharp violent motion. Heavily motion-blurred, dark, no clean read, disorienting. Rendered with clean per-object motion blur rather than optical smear. Lighting: near-black, harsh partial specular highlights only, deep ambient occlusion. Audio: torn shards of noise cutting hard to silence at 1.5s. 1.5s HARD CUT 1.5s–6.5s — CLINIC WIDE (establishing, three zones) First frame already holds all three: @image_1 FG left third (x20%, y60%) angled to his console; @image_3 MG center (x50%, y50%) upright, square to camera, thin electrode pads on both temples, eyes open and completely blank with a clean circular catchlight in each; @image_4 FG right third (x80%, y62%) seated in the waiting chair in his black suit, eyes on @image_3. No empty frame. Lens lock: 84° diagonal field of view, wide-angle character, virtual camera at the far wall roughly 4 to 5 meters back so all three zones read across full width; rectilinear, no fisheye, deep even focus with no corner falloff. Camera: locked, eye-level, extremely slow imperceptible push-in on smooth interpolated motion. Action: near-static; @image_3 one slow blink around 4.0s; @image_1 minimal hands on console; @image_4 still. Lighting: bright cold clinical high-key, even blue-white ceiling wash at roughly 5600K, faces fully readable, no warm fill. Screen-space reflections of the ceiling panels and the three figures across the polished floor. Soft ambient occlusion under every chair leg, cable and equipment edge. Audio: clinic room tone, steady slow monitor beep. 6.5s HARD CUT 6.5s–8.0s — INSERT: THE CHOICE Tight insert on the extraction console screen. A clean minimal game-style UI on a pale blue-white field reads exactly "Delete this memory?" with "Yes" and "No" below it as two thin rounded-outline buttons; a soft rectangular cursor highlight rests BETWEEN them, touching neither, indicating an undecided state. All three text strings rendered exactly as written, correctly spelled, cleanly kerned and perfectly sharp. Nothing else on the screen — no icons, no window chrome, no numerals, no patient data, no other text. Lens lock: 29° detail, camera close on the screen, screen critically sharp, room falling into clean computed bokeh behind. Lighting: screen self-lit and emissive, cool, casting a soft blue spill onto the console housing with visible fingerprint smudging on the glass and gentle bloom at the screen edges. Audio: one short interface beep, then quiet. 8.0s HARD CUT 8.0s–9.5s — INSERT: KAI'S GLANCE Tight emotional close-up on @image_1's face. Without turning his head, his eyes flick to screen-right toward @image_4, then straight back down to the screen. Face controlled, but the glance carries suspicion. Lens lock: 18° classic telephoto, tight, background compressed and cleanly defocused. Lighting: cool emissive machine glow from the screen side, the other side falling to soft shadow with ambient occlusion in the eye sockets and under the jaw, both eyes catching a crisp circular highlight. Audio: clinic silence, faint beep. 9.5s HARD CUT 9.5s–19.0s — WARM MEMORY MONTAGE (four quick shots) 9.5s–11.5s: @image_5 (child Leta) alone in a playground sandbox, small in a warm open frame. Lens 47° natural. Warm golden afternoon light with soft volumetric haze, everything cleanly resolved to the far fence line. 11.5s HARD CUT, 11.5s–13.5s: @image_6 (teen Damen) at a back-row school desk, turns to camera, smiles, looks away, then smirks and spins a pen in his fingers. Windows behind throw strong rim backlight with visible volumetric shafts. Lens 29° portrait toward camera. Warm classroom light, backlit haze, gentle bloom on the window edge. 13.5s HARD CUT, 13.5s–16.0s: @image_2 (youthful Leta) laughing among friends in a bar. Lens 47°. Warm amber, the crowd behind her rendered as individually varied figures falling into clean circular bokeh, never flat repeated copies. Bloom on the practical lights. 16.0s HARD CUT, 16.0s–19.0s: @image_2 on a beach at sunset, water gold; she enters from mid-left walking in, friends run toward the water in the background as dark blown-out silhouettes seen from behind. Lens 84° environmental. Golden backlight, glowing water with screen-space reflections and specular sun glitter, gentle bloom on the horizon. Audio across montage: warm overlapping laughter and chatter, cutting hard to full silence exactly on the last beach frame. 19.0s HARD CUT 19.0s–21.0s — CLINIC CLOSE (Kai softens) Close-up on @image_1. The corner of his mouth lifts a fraction; his eyes warm for a beat. Because his face carries no habitual lines, this micro-expression reads as an unfamiliar movement on an unused face. Lens lock: 18° tight. Lighting: cold clinical, unchanged from earlier clinic beats. Audio: the abrupt silence left by the cut laughter, then the slow beep returns. 21.0s HARD CUT 21.0s–30.0s — CLINIC MEDIUM (the exchange) Two-person medium. @image_4 crosses from the right waiting zone in his black suit and his hand lands on @image_1's shoulder from behind, the suit wool creasing with clean simulated cloth folds. @image_4 speaks: "How much longer." @image_1's face hardens instantly; he answers level, without turning far: "Wait in your chair." @image_4 holds the look a beat too long, then turns and walks back to the waiting chair. End on @image_1's hardened profile. Lens lock: 47° standard normal, camera 3 to 4 meters, both readable. Camera: locked, slight settle on smooth interpolated motion. Lighting: cold clinical high-key, both faces readable, no warm fill, ambient occlusion where his hand meets the coat shoulder. Audio: @image_4's line begins within 0.3s of this shot; clean close English dialogue over the silent room and the slow beep; lips still except on the two lines. POSITIVE LOCKS Every beat is rendered in the high-end real-time game cinematic style defined above, with no exceptions. All clinic beats share one cold blue-white clinical grade and the same three-zone geometry; memory beats are warm and nostalgic; the teaser is near-black and fragmented. Faces come 100% from the references and never change between cuts. Kai holds screen-left in his white clinical coat, Damen holds screen-right in his black suit, in every clinic beat. Leta's light chestnut hair and golden-amber eyes are identical in @image_2, @image_3 and @image_5. Only the two quoted English lines are spoken; everyone else stays silent with still lips. Every transition is a hard cut except the specified slow push-ins. No gore, no blood, no visible violence beyond the abstract teaser flashes.

Part 2: the reveal (0:30 to 1:00)

What it does: generates the second half in one pass, first love, the fight, the family beats, the accusation, the backyard reveal and the standoff, ending on the title card. The upload order changes here: 1 Kai, 2 Leta at 29, 3 Damen, 4 Leta at sixteen, 5 Damen at sixteen, 6 Leta's parents.

SCENE CONTEXT Closing half of the trailer, rendered as a HIGH-END REAL-TIME GAME CINEMATIC. Warm first love cracks into a violent marriage; the operator starts to feel what he is erasing; buried family secrets surface across a cold table, a slammed door, and a scream at the parents. A memory of four people digging at dusk reveals the operator's own younger face reflected in the woman's terrified eyes. In the clinic he recoils and faces the husband as a single tear runs down her blank face, then the image dissolves to the title. RENDERING STYLE — MANDATORY, APPLIES TO EVERY BEAT: this must read as a top-tier real-time rendered game cinematic in the register of a modern narrative console title — near-photoreal fidelity but unmistakably a rendered engine image rather than photographed footage. Every tell below must be present in every shot. SKIN: full micro-detail — pores, fine lines, moles and stubble all rendered — but with the characteristic engine subsurface look: smooth even translucency beneath the surface and a LIGHT UNIFORM SPECULAR SHEEN across the forehead, the bridge of the nose, the cheekbones and the chin, as though very faintly damp. Slightly waxy, slightly too clean, with no oil variation and no blotchiness. Detail evenly distributed across the whole face with no falloff. EYES: glassy and high-specular, with a clean bright sclera, a crisp perfectly circular catchlight in each eye, a strong wet highlight on the cornea and a sharply defined limbal ring. Slightly too perfect, slightly synthetic. Individually rendered eyelashes. HAIR: rendered in defined strand clumps and card-like groupings rather than fully individual chaotic strands, with a uniform anisotropic sheen running along the length and clean silhouette edges. MATERIALS: physically based shading throughout, every surface with correct roughness and reflectance, visible screen-space reflections on glass, metal, polished floors, water and screens, and soft ambient occlusion darkening every crease, seam and contact point. All fabric reads as simulated cloth — clean folds, slightly stiff drape, uniform wrinkle logic. IMAGE AND POST: everything critically sharp across the full frame with no lens softness, no field curvature and no corner falloff. Detail and contrast do NOT diminish with distance — backgrounds are as cleanly resolved as foregrounds. Where depth of field is used it is computed and clean with perfectly circular uniform bokeh. Motion blur is present and correct but clean and per-object rather than optical smear. Gentle uniform bloom on the brightest highlights. Subtle chromatic aberration applied evenly as a post effect. A fine even film grain overlay laid uniformly across the whole image. Slightly elevated contrast and clean highlight rolloff. All camera movement is smooth interpolated virtual-camera motion with no digital jitter, except where handheld micro-motion is explicitly specified. FORBIDDEN: no illustration, no anime, no cel shading, no stylised or exaggerated proportions, no painterly look, no concept art, no cartoon, no low-poly, no visible polygon edges, no texture seams, no flat shading, no clay render, no grey-boxed or unfinished surfaces, no flat repeated crowd instances. Fidelity must be maximal and the image must look expensive — only the LOOK is rendered rather than photographic. ASSET REGISTRY (reference handles = upload order; identity is 100% from each reference) @image_1 — KAI WEIL, 30. Operator. Fair skin with a cool under-slept pallor, dark brown almost black hair cut short and neatly, dark brown eyes, lean and narrow through the shoulders with visible tendon at the wrists and neck. HIS FACE CARRIES ALMOST NO MIMIC LINES — no crow's feet, no forehead lines, no nasolabial folds, no frown lines — a fully neutral trained mask, strangely smooth for his age. A thin pale horizontal scar, 15mm, at the left temple just inside the hairline. Wears a clean WHITE clinical coat, open, over a dark charcoal-grey crew-neck base layer and slim dark charcoal trousers. A thin matte dark grey interface cuff on the right wrist. 100% matches the reference. He ALSO appears younger inside the backyard memory — the same face made visibly younger, with the temple scar absent. @image_2 — LETA MOREN, present. 29. LIGHT CHESTNUT hair falling loose past the shoulders in a soft wave with loose face-framing strands, GOLDEN-AMBER eyes with visible gold flecks, faint freckling across the nose bridge, a small dark mole below the outer corner of the left eye, soft natural makeup, small gold studs, a slim gold band on her finger. Softer, melancholic. HER FACE, HAIR COLOUR, EYE COLOUR, FRECKLING, MOLE AND JEWELLERY ARE 100% FROM THE REFERENCE AND NEVER CHANGE. HER CLOTHING IS DIFFERENT IN EVERY SCENE AND IS SPECIFIED PER BEAT BELOW — the clothing in the reference image is NOT used anywhere in this film. @image_3 — DAMEN VOSS, 32. The husband and antagonist. Tall and solidly built through the chest with the beginnings of softness at the waist. ASH-BLOND HAIR WITH CLEAR GREY AT BOTH TEMPLES, combed back with visible product. Fair skin with a faint flush and broken capillaries across the nose and cheeks. PALE GREY-BLUE cold eyes beneath a heavy brow. HIS FACE IS CARVED BY FROWNING — two deep permanent horizontal lines across the forehead and a third shorter one between the brows, present even at rest, plus pronounced nasolabial folds. Wears an EXPENSIVE BLACK WOOL SUIT, jacket open and unbuttoned, over a white dress shirt with the top two buttons undone and NO TIE. A heavy steel watch, a plain gold wedding band. Cold and controlled, with nervous slightly mechanical smiles. 100% matches the reference, and his suit stays the same in every adult beat. @image_4 — LETA as a teenager, about 16. Built from @image_2's features: light chestnut hair, golden-amber eyes, the same freckling and the same mole below the left eye, softer and younger. Wears a plain pale blue school shirt with the sleeves pushed up and a plain dark skirt. First-love classroom only. @image_5 — DAMEN as a teenager, about 16. Built from @image_3's features but softer and boyish: ash-blond hair, pale grey-blue eyes, no grey and no frown lines yet, in a plain dark school jumper over a pale shirt. First-love classroom only. @image_6 — LETA'S TWO PARENTS, an older couple roughly 55 to 65, plainly and expensively dressed in muted neutral tones. The mother carries a clear echo of @image_2's features. Table and study only. LETA'S WARDROBE — DIFFERENT IN EVERY SCENE, MANDATORY: her face and hair never change, but she is dressed differently in every beat and the costume must be unmistakably distinct each time. Under no circumstances does she wear the same outfit in two different locations. CLINIC — a plain pale grey-blue hospital gown in soft matte cotton, loose at the shoulders, plain round neckline, no fastenings visible, no pattern, no markings. Two thin flat circular electrode pads in matte dark grey polymer against both temples with slim pale grey cables running back out of frame, her hair tucked so both pads are clearly visible. All jewellery removed except the slim gold band. Barefoot or in plain pale slippers. THE FIGHT, bedroom at night — a fine ivory silk camisole with narrow straps and loose matching ivory silk trousers, the fabric catching the cold lamp light with a soft sheen. Hair loose and slightly disordered. She is undressed for bed and physically unguarded, which is the point. THE TABLE, formal family dinner — a structured DEEP FOREST-GREEN long-sleeved dress in matte crepe, high closed neckline, buttoned or fastened all the way to the throat, fitted and formal. Hair pulled back neatly. She is fully covered and immaculate, and it reads as armour. THE DOOR, ordinary daytime at home — a soft PALE CREAM chunky-knit jumper, slightly oversized, over plain mid-blue jeans, sleeves pushed back. Hair loose. Domestic, off guard, caught unprepared. THE ACCUSATION, parents' study — a long CHARCOAL-GREY wool overcoat worn OPEN and still on indoors over a plain black fine-knit top and dark trousers, a dark scarf loose around her neck. She has come in from outside and has not taken her coat off, because she came to demand an answer. Hair windblown. THE BACKYARD, dusk — THE SAME CHARCOAL-GREY OVERCOAT, black knit and dark trousers as the study beat, but now dishevelled and marked: the coat hem and one sleeve smeared with earth, a scuff of soil at one knee, the scarf pulled loose and hanging, hair pulled out of place. This deliberate continuity links the two memories into one evening and must be preserved exactly. GLOBAL LOCKS (hold across every beat) Identity is 100% from the references, the same faces as Part 1, never redesigned. Clinic beats reuse the exact cold clinical high-key grade from Part 1 so both halves read as one film. Memory beats carry their own scene-appropriate light: warm for young love, cold moonlight for the fight, cold institutional for the table, fading dusk for the backyard. Only the scripted English scream line is spoken; all other lips stay still; no narration, no offscreen extra voices. The slow monitor beep runs under clinic beats. NO TEXT anywhere except the final title card — no logos, no signage, no name badges, no ID cards, no lanyards, no wall labels, no equipment markings, no prints or graphics on any garment. SHOT SEQUENCE — controlled multi-shot, real-time, hard cuts only unless stated. 0.0s–1.5s — MEMORY: FIRST LOOK @image_4 in her pale blue school shirt and @image_5 in his dark school jumper, in the same classroom, younger, exchanging a shy glance of first love. Lens lock: 29° portrait. Lighting: warm, soft, nostalgic, with gentle bloom on a bright window edge and soft ambient occlusion under the desks. Audio: warm room tone building. 1.5s HARD CUT 1.5s–4.0s — MEMORY: THE FIGHT A married bedroom at night. A bedside lamp is knocked askew, throwing hard cold light. @image_2 IN THE IVORY SILK CAMISOLE AND LOOSE SILK TROUSERS, hair loose and disordered, and @image_3 in his black suit are mid-argument, bodies tense and close. Something shatters just offscreen. Lens lock: 84° close environmental, unsettled handheld micro-motion (operator breath and weight shift, no digital jitter). Lighting: cold hard moonlight, low-key, deep blue shadows with strong ambient occlusion, the askew lamp emissive and casting a hard dynamic real-time shadow, the silk catching a cold specular sheen. Audio: the warm bed of the prior beat drops away as raised voices rise, one object smashes offscreen, then hard cut to full silence. 4.0s HARD CUT 4.0s–6.0s — CLINIC: KAI FEELS IT @image_1 frowns, real emotion breaking through for the first time — and because his face carries no habitual lines, this reads as an unfamiliar movement on an unused face. In the right waiting zone, @image_3 in his black suit tenses and starts to rise from the chair. Lens lock: 47° medium. Lighting: cold clinical high-key at roughly 5600K, faces fully readable, no warm fill, screen-space reflections across the polished floor. Audio: silence, then the beep, faintly uneven. 6.0s HARD CUT 6.0s–8.0s — MEMORY: THE TABLE A long family dining table. @image_2 IN THE DEEP FOREST-GREEN HIGH-NECKED DRESS, hair pulled back neatly, sits with @image_6, her parents; all three silent, seated far apart, cold distance between them. Lens lock: 84° wide down the length of the table. Lighting: cold, even, flat and institutional, with soft ambient occlusion at every chair and place setting and the table surface carrying clean screen-space reflections. The green crepe reads matte against the cold light. Audio: heavy formal silence, faint room tone. 8.0s HARD CUT 8.0s–10.5s — MEMORY: THE DOOR @image_3 stands in a doorway, smiling at someone offscreen; he notices @image_2 — IN THE PALE CREAM OVERSIZED KNIT AND MID-BLUE JEANS, hair loose — his face changes, he steps to the door and slams it hard directly into the lens so the door fills the frame. Lens lock: 47° widening toward 84° as the door looms into camera. Lighting: neutral daytime interior falling dark as the door closes, the light pinching out along the closing edge. Audio: the slam is loud and hard, and it bleeds across the next cut. 10.5s HARD CUT (slam carries over) 10.5s–12.0s — CLINIC INSERT: THE HAND Tight insert on @image_2's hand resting on the chair, the pale grey-blue gown sleeve at the edge of frame and the slim gold band catching a crisp specular; her fingers twitch once. Lens lock: 29° detail, background falling into clean computed bokeh. Lighting: cold clinical. Audio: the carried slam lands here; the monitor beep stutters for a fraction of a second. 12.0s HARD CUT 12.0s–15.5s — MEMORY: THE ACCUSATION A parents' study. @image_2 IN THE CHARCOAL OVERCOAT WORN OPEN OVER A BLACK KNIT, dark scarf loose, hair windblown, still in her outdoor coat indoors, screams at @image_6: "You knew. You both knew." Then drops to her knees, breaking into tears, the heavy coat folding around her as she sinks. Lens lock: 47° medium, following her down as she sinks on smooth interpolated motion. Lighting: tense, harder contrast, cool, with deep ambient occlusion in the room's corners. Audio: rising tension into the scream, line begins within 0.3s of this shot, then raw crying. 15.5s HARD CUT 15.5s–20.0s — MEMORY: THE BACKYARD (the reveal) Dusk. @image_2 IN THE SAME CHARCOAL OVERCOAT, BLACK KNIT AND DARK TROUSERS AS THE PREVIOUS BEAT BUT NOW DISHEVELLED AND EARTH-MARKED — the coat hem and one sleeve smeared with soil, a scuff of earth at one knee, the scarf hanging loose, hair pulled out of place — hides behind a tree trunk in a family backyard, peeking out; through the gaps between trees, four people are digging a hole with a mound of fresh earth beside it. She is horrified. Someone grabs her from behind and clamps a hand over her mouth so she cannot scream. From around 18.5s, a slow deliberate push into a macro close-up of her wide terrified eye: in the wet reflection on the eye, @image_1's face, visibly younger than present day, surfaces clearly for a held fraction of a second before the beat ends. Give this reflection real screen time and a clean bright specular so the face reads unmistakably; do not rely on a single frame. Lens lock: opens at 18° telephoto through the trees with blurred tree trunks occluding the lower foreground; ends on extreme macro on the eye. Lighting: fading dusk, low warm-to-blue, tree-filtered with soft volumetric haze between the trunks; the four diggers read as dark silhouettes with no facial detail; on the macro, a clean circular catchlight in the eye carries the reflected younger face with correct specular reflection on the wet cornea. Audio: her heavy breathing, distant shovel strikes and shifting soil, breathing cut off hard as the hand covers her mouth. 20.0s HARD CUT 20.0s–26.0s — CLINIC: THE STANDOFF (climax) @image_1 recoils hard back from the console, shaken, steps back and turns, and @image_3 in his black suit is standing right there, close, face angry and tense with the frown lines deepened, the look of a man about to be exposed. They face each other, no dialogue. Between and behind them, in the chair in deep mid-ground, @image_2 IN THE PALE GREY-BLUE HOSPITAL GOWN WITH THE TEMPLE ELECTRODES sits with open blank eyes as a single tear runs down her cheek. Lens lock: 47° two-shot, camera holding both men with @image_2 readable between them in the background, everything cleanly resolved to the back wall. Camera: locked, a slow tightening push on smooth interpolated motion. Lighting: cold clinical high-key, all faces readable, a clean bright specular on the running tear, ambient occlusion where the two men's shapes overlap. Audio: the same slow monitor beep, otherwise silence. 26.0s–30.0s — DISSOLVE TO TITLE The slow push continues onto @image_2's tear; the tear becomes a small ripple of water and the whole image dissolves into that ripple, the one soft transition in the piece. The water carries clean screen-space reflections and gentle caustic light. The title resolves out of the water, reading exactly: LETA in a thin light sans-serif in pale cool grey, correctly spelled, cleanly kerned and perfectly sharp, centred on a clean dark field. Then a single clean call-to-action card follows on the same dark field. NO other text, no numerals, no dates, no credits, no logos. Lighting: water reflection, minimal, resolving to a clean title field with a faint even bloom. Audio: silence gives way to a single low sustained tone that fades to full silence. POSITIVE LOCKS Every beat is rendered in the high-end real-time game cinematic style defined above, with no exceptions. The same faces as Part 1, 100% from the references, never redesigned; @image_1's backyard reflection is the same face made visibly younger. LETA'S FACE, HAIR AND EYES ARE IDENTICAL IN EVERY BEAT WHILE HER CLOTHING IS DIFFERENT IN EVERY SCENE EXACTLY AS SPECIFIED — hospital gown in the clinic, ivory silk in the fight, forest-green high-necked dress at the table, pale cream knit and jeans at the door, charcoal overcoat in the study and the same coat earth-marked in the backyard. She never wears the same outfit in two different locations. Damen's black suit, ash-blond hair with grey temples and permanent frown lines are consistent in every adult beat; his teenage version has neither the grey nor the lines. All clinic beats match Part 1's cold clinical grade and geometry; each memory keeps its own scene-appropriate light. Kai holds screen-left in his white clinical coat, Damen holds screen-right in his black suit, in every clinic beat. Only the English scream line is spoken; all other lips stay still. Every transition is a hard cut except the final dissolve into the tear-ripple and the two specified slow push-ins. No gore, no blood, no bodies, no visible violence beyond the offscreen smash and the covered mouth.

What else can these same assets do?

The reference library was already built once for the trailer, and the same characters and locations move directly into Marketing Studio for a batch of ad creatives, no rebuilding required. Same reference set, so every version still reads as the same game.

That's what an AI-native creative suite buys here: the trailer and the ad batch come out of one pipeline, not two separate projects built from scratch.

How do you enter the Higgsfield Global Film Festival?

Pro tip: if you want to go deeper into AI filmmaking, start with Higgsfield Academy. The AI Filmmaking Pipeline course walks the same method as this guide, scripting, asset creation and scene-by-scene prompt engineering, taught on the football drama from How To Direct an AI Short Film Like a Filmmaker. Learn the pipeline on a trailer, scale it to a short film, then take that film somewhere.

Higgsfield is putting real money behind exactly that: the Higgsfield Global Film Festival opens on August 7, 2026, with a $1,000,000 prize pool. Entry is open to anyone 18 or older with an active subscription, solo or in a team of up to four. Films can be any genre, minimum 3 minutes long, generated entirely inside higgsfield.ai's own Generation Window between the opening date and the deadline, nothing brought in from outside. Submissions close August 31, 2026, at 23:59 PT. Prizes: $500,000 for first place, $200,000 for second, $100,000 for third, a $100,000 Audience Choice award, and ten additional $10,000 prizes. Winners are announced in early October 2026. Filmmakers keep full rights to their own film, Higgsfield only receives a promotional license. Full rules are on the festival page.

Not sure what a finished piece looks like at this level? The Explore page is full of showcase examples made by the community and the Higgsfield team, trailers, shorts and full scenes, most with visible prompts you can pull apart and learn from. A one-minute trailer like the one above wouldn't meet the 3-minute minimum on its own, but everything in this guide, the locked characters, the locations, the scene-by-scene workflow, scales directly into a festival submission.

How To Make a Game Trailer with AI (Full Workflow + Prompts)

Try Cinema Studio

Got any questions left?

In general, you need a way to generate consistent character and location images, then a way to turn them into video. Inside Higgsfield specifically, that means Soul Cinema for the still references and Seedance 2.5 720p for every video shot, all inside Cinema Studio.

You lock the face as a reference sheet before generating a single scene, then attach that same file to every shot the character appears in. In Higgsfield the prompt reads references by upload position, so keep the attachment order identical every time, and write the same physical details into the prompt as a backup.

Count it in three stages rather than estimating the trailer as a whole: one generation pass, plus retries, for each character reference, one for each location, and one for each video pass. In a workflow like this the images are quick and cheap, and almost all of the time and cost sits in the video passes.

Yes. In Higgsfield, once a character or location is locked, it drops directly into Marketing Studio to batch short ad variants, no separate shoot required.

Some video models, including Seedance, don't process negative instructions the way text models do. Writing "she doesn't smile" can produce the opposite result, so it's safer to describe the state you actually want.

A cinematic trailer dramatizes character and world, while gameplay footage shows how the game actually plays. Some studios capture cinematic trailers directly in their game engine, while an asset-first AI workflow builds them from locked character and location references animated shot by shot instead.

by Higgsfield

Share article

Discover more

View all