Blog

Common AI Video Prompt Mistakes (And the Fixes That Actually Work)

HiggsfieldAug 3, 202612 min
Common AI Video Prompt Mistakes (And the Fixes That Actually Work)

Most AI video prompt mistakes fall into three groups. Describing a subject too loosely lets faces and clothing shift on every run. Leaving the camera, light, or movement undefined makes atmosphere and motion come out wrong. Skipping the blocking on a multi-scene shot lets characters drift out of position. Each gap gets filled by the model on its own.

Why does the same prompt give a different result every time?

A prompt only carries what gets written into it. Everything left out, an exact age, a light source, the end point of a camera move, still gets decided, just by the model instead of by you. That's why the same wording gives a slightly different result on every run. The model is closing whatever the prompt left open.

A reference image fixes appearance. It shows what a face or a room looks like, but says nothing about how a hand moves, where the light falls a second later, or where the camera ends up by the last frame. Without a written brief covering that half of the shot, the model is guessing at it regardless of how strong the reference is.

Detail competes for the model's attention. A single dense paragraph trying to cover the face, the gesture, the lighting, and the camera all at once forces the model to split its attention, and hands, expressions, and exact positioning are usually what suffers for it.

Every extra person and every extra beat multiplies this. A single person in a static shot only has to hold up under one interpretation. Five people, a moving camera, and several lines of dialogue is five or six interpretations running at once, and one of them drifting is enough to break the scene.


What are the most common AI video prompt mistakes?

What are the most common AI video prompt mistakes?

Mistake

Why It Happens

The Fix

Characters look plastic or fake

Described only in general terms, no reference or distinguishing detail

Exact age, build, and clothing, backed by a reference or a trained Soul ID

Extra fingers, fused hands, AI slop

Points of contact were never specified, so the model fills them in on its own

Name where every object sits and exactly how hands make contact

Wrong emotion, unclear what's happening

Only the general idea of the scene was written, not a visible expression

Tie the emotion to a specific action: a gaze, a held breath, a sigh

Atmosphere and lighting miss the mood

Light source, angle, and time of day never named

State them directly, or set genre and lighting as parameters in Cinema Studio

Unnatural movement

Motion described as one adjective instead of an actual path

Spell out the turn, tilt, and angle, or lock it as a Cinema Studio setting

Camera clutter, lost focal point

Shot list names angles but not what each shot is for

Give every shot one job and cut what doesn't serve it

Characters lose their positions in complex scenes

Only the overall idea is written, not who stands where

Storyboard the scene in Popcorn before generating the full sequence


How do I write a strong AI video prompt?

A strong prompt is closer to a shot list a crew could work from than to a description. Follow these in order and most of the seven mistakes below never get a chance to happen.

Step 1: Set the scene and lock the cast. Name the setting, the duration, and whether it's meant to feel real-time or stylized. Then give each person an exact age, build, and outfit, plus one distinguishing detail, and say plainly that these stay identical across every cut. For a face that needs to reappear in other videos, train it once in Soul ID instead of redescribing it here.

Step 2: Block the scene before writing any action. Say where each person or object sits relative to the others and to the camera, and which direction they move. Keep that direction and the camera's side consistent through the whole scene, so nothing flips sides partway through.

Step 3: Define the camera as a physical setup, not a mood. Field of view, distance from the subject, and whether the camera is static, handheld, or moving, with the movement type named directly rather than described as "dynamic." Where it's available, set this directly in Cinema Studio instead of writing it out.

Step 4: Write the action as a shot list, one beat at a time. Give each beat one clear action, described as a start position and an end position rather than a mood word. In the same pass, name every point of contact, a hand on a shoulder, fingers around an object, since this is exactly where extra fingers or fused hands come from if it's left unsaid.

Step 5: Set physics and lighting as real events. Say how things should behave under weight and gravity, fabric, fire, water, a body catching itself, so movement reads as physical rather than simulated. Light gets the same treatment: one named source, one direction, one color, with anything that shouldn't double up, a second sun, a stray flare, ruled out directly.

Step 6: Decide the audio and the finish. State what's actually heard, whose line belongs to whom, and whether music plays at all. Close with the visual reference, a film stock, a grain level, a color grade, and rule out what shouldn't appear: no text, no logos, no watermarks.

For anything with more than one character or more than one beat, this same information still applies, it's just worth planning as a storyboard before the full sequence gets generated, rather than writing all of it into one long paragraph.


Here's where that plays out in practice. The seven mistakes below fall into those same three groups: the first three come from an underdescribed subject, the next three from an undefined camera, light, or movement, and the last from a scene that was never blocked. Each one below has the prompt that caused it and the rewrite that fixes it.

Why do my AI characters look plastic or fake?

A character described only in general terms gets reinvented on every generation. "A group of friends around a fire" leaves age, posture, and face entirely open, and the model answers that differently each time it runs. The plastic look comes from exactly that, nobody told the model who these people are in enough detail to recognize them a second time.

Bad prompt: five friends at a campfire, cast left open, no assets attached.

Group of friends sitting around a campfire at night talking about meeting more often. Very cinematic, emotional, 4k, hyperrealistic, best quality. A man in the middle and some girls around him in scarves and jackets. Add more people if it looks better. 0:00-0:03 wide shot, man says "I don't know why we let a whole year go by". 0:03-0:06 close up of the girl "Then let's stop letting it". 0:06-0:09 another angle, wider. 0:09-0:12 close up of the girl again, smiling. 0:12-0:15 back to wide shot. Dynamic cinematic angles, pick the best composition for each cut. Flickering firelight effect on the faces, teal and orange grading. Big beautiful flames in the foreground. Warm nostalgic friendship mood.

In the campfire scene, the prompt left the cast open with "add more people if it looks better" and attached no assets, and the faces came back interchangeable.

Good prompt: five named characters, fixed blocking, locked light direction.

SCENE CONTEXT: Five friends around a small campfire on a grassy coastal bluff, late blue-hour dusk. They agree to meet more often. Real-time, about 15s, five segments, four hard cuts. CAST, five distinct people, no repeated face or wardrobe: P1 man 32, dark-blond, light grey henley. P2 man 29, dark curls, stubble, dark green shirt-jacket. P3 woman 27, long straight black hair, olive field jacket. P4 woman 30, wavy copper hair, cream cable knit. P5 woman 26, ash-blonde low bun, navy cardigan. Exactly five people. No twins, no duplicated wardrobe, no extra figures. BLOCKING: Fire foreground center, x50 y74, blurred. Group in a semicircle beyond it, 1.5 m from flames. Camera stays on one side of the fire; the 180 axis runs through the fire and is never crossed. P3 and P4 stay screen-left looking camera-right; P1, P2, P5 stay screen-right looking camera-left. First frame already shows three people seated in position, no empty establishing frame. OPTICS AND CAMERA: Seg 1 and 5, 47 degree FOV, camera 3.5 m, group two/three-shot. Seg 2 and 4, 29 degree FOV, camera 4.5 m, close-up with a blurred listener's shoulder in the near foreground. Seg 3, 47 degree FOV, camera 3 m, medium close-up. No lens drift inside a segment. Shoulder-held at seated eye level, slow operator breath, shutter 270-360 degrees, natural motion blur. No zoom, no push-in, no digital jitter. BEATS: 1. Group wide. P1 to P3: "I don't know why we let a whole year go by." Hard cut. 2. P3 close-up, eyes on P1 off-frame right: "Then let's stop letting it." Hard cut. 3. P4 medium close-up, mug in hands: "Same fire, same people. Once a month." Hard cut. 4. P5 close-up, three-quarter left, sparks crossing frame: "I'm holding all of you to that." Hard cut. 5. Back to P1 with P2 behind: "Then it's a plan." Long held static final beat. Only the named speaker talks; listeners keep lips closed. No offscreen voices. PHYSICS: Fire burns like real wood, fast irregular tongues, tips breaking off, uneven height, thin smoke drifting one way. Sparks rise on turbulent curved paths and fade, never in straight lines. Light sea wind moves loose hair and smoke consistently. Bodies carry weight, mugs tilt with liquid inertia. LIGHTING: Key is the fire itself, 2200-2700K, from below. Fill and rim from cold 6500K dusk sky, separating each person from the dark ridge. Firelight swells slowly and organically with the flame shape, shadows rolling off gradually. No high-frequency strobing, no stepped luminance jumps, no unmotivated source. Warm catchlight in every eye. AUDIO: Only the five scripted English lines, close and dry. Continuous wood crackle, quieter on close-ups, soft wind and distant surf. Ambient ducks under speech. No music, no subtitles. LOOK: Photoreal, shot on ARRI Alexa 35, natural grain, warm fire against cold blue dusk, visible pores, freckles, fabric weave. Same faces, wardrobe, seating order, fire size and light direction across every cut. No text, no logos, no watermarks.

Write each person as a specific age, build, and outfit, backed by a reference image or asset. Then write in the small things people do without thinking: a head turn, a pause before a line, a sigh. For a face that needs to reappear across other videos, Soul ID trains it once from a set of reference photos and holds it from there.

Why does AI video generate extra fingers, two suns, and other slop?

AI slop covers a range of things that shouldn't be physically possible, a third arm, a second sun, a sixth finger, and all of it comes from the same place: something in the frame was never described, so the model filled it in on its own. In the couple's living room scene, the prompt asked for "gorgeous sunlight bursting through the curtains" with "pulsing god rays" and never pinned that down to one fixed source, and what came back showed two suns. The good version named one source with a warmth and a direction, and got one sun. Hands work the same way: an unnamed position is usually where an extra finger or a fused hand comes from.

Bad prompt: a slow dance with the light source and the hands both left unspecified.

SCENE: A young couple slow dances in the living room of a 1950s home, late afternoon. One continuous take, about 8 seconds, cinematic and dreamy. SUBJECTS: A romantic vintage couple, he in a shirt and suspenders, she in a pretty floral dress. Beautiful faces, 8k, hyperrealistic, best quality. CAMERA: Camera pulls back from an extreme close-up of their faces into a wide shot, dropping lower as it goes for a more dramatic angle. Find the best composition while retreating. Cinematic dolly zoom feel, epic reveal of the room. ACTION: 0:00-0:03 close up of their faces touching. 0:03-0:05 they sway gently in place, barely moving. 0:05-0:08 wide shot of the whole room as they keep swaying. HANDS AND CONTACT: Their hands come together into one soft romantic embrace, fingers intertwined and melting into each other, hands blending as they hold each other tight. SET: Cozy retro living room packed with vintage props, lamps, pictures, furniture everywhere, lots of nostalgic detail in the background. LIGHTING: Gorgeous sunlight bursting through the curtains with cinematic lens flare and pulsing god rays, dynamic light, magical atmosphere, glowing highlights. AUDIO: Romantic vintage vibe, humming, soft laugh. LOOK: Cinematic, dreamy, nostalgic, award winning, 8k, hyperrealistic.

Good prompt: one named light source, every point of contact described finger by finger.

SCENE: A young couple slow dances in the living room of a 1950s home, late afternoon. One continuous take, about 8 seconds, real-time. SUBJECTS: Man 30, dark combed-back hair, white shirt with suspenders and dark trousers. Woman 27, soft waves, floral tea dress, low heels. Both faces stay identical from the extreme close-up to the wide shot. CAMERA: Single continuous crane back at a constant eye level, 1.5 m above the floor. Opens on an extreme close-up of their touching foreheads, glides steadily backward and slightly up into a medium wide shot, holding the couple centered in the middle third with clear headroom. 40 degree field of view, shutter 270-360 degrees, no zoom, no low-angle tilt, no speed change. ACTION: They rotate slowly and continuously through a full quarter turn as the camera retreats. Each step lands with a real heel-to-toe contact and weight transfer, soles gripping the rug, no sliding, no floating. She laughs softly midway and lets her head tilt against his shoulder. HANDS AND CONTACT: His right hand rests flat and open on her upper back, fingers distinct. Her left hand sits on his shoulder with five separate visible fingers. Their other two hands are clasped at chest height, two anatomically separate hands with individual fingers, knuckles and skin folds readable throughout the turn. Hands never fuse, never sink into fabric. SET: Period living room, sofa, patterned rug, table lamp, framed picture, wooden radio cabinet, three curtained windows behind the couple. All furniture and props keep rigid, stable geometry while the camera moves; nothing warps or shifts scale with parallax. LIGHTING: Warm 4000K afternoon sun as backlight through sheer curtains, constant intensity, soft rim on hair and shoulders. Gentle even fill from camera-left. Volumetric beams and dust motes stay steady with slow natural drift. No lens flare, no pulsing rays, no flicker. AUDIO: He hums a slow waltz melody under his breath, she laughs quietly once. Faint fabric rustle and shoe contact on the rug. Room tone only, no music track. LOOK: Photoreal vintage cinema, warm muted pastel grade, soft contrast, fine grain, readable skin pores and fabric weave. No text, no logos, no watermarks.

Describe every object in the frame and exactly where it sits, the light source included, down to where the main point of light actually is. Left undescribed, the model has to invent the scene itself, and that's where the duplicates and extra limbs come from. The same care covers hands and small points of contact, how many fingers, where one rests, what stays separate from what. If the scene is complex enough that several of these need to hold at once, storyboarding it in Popcorn first means the sequence gets animated from a prepared plan rather than generated at random.

Why doesn't the emotion in my AI video read?

A mood word doesn't tell the model what to put on a face. "Sad," "tense," "joyful" leave the actual expression open, and most generations settle on something flat and pleasant that misses the scene. A scene can have a real idea behind it and still read as confusing, simply because the face was never told what to do.

Bad prompt: the emotion left as mood words.

SCENE: A beautiful young woman lying on a red carpet in a dark room, breathing heavily. Cinematic, moody, about 7 seconds, dramatic. SUBJECT: Gorgeous woman with wavy hair, cream top, colorful silk scarf, big gold hoops and shiny gold bracelets. Stunning, 8k, hyperrealistic, best quality. CAMERA: 0:00-0:03 camera slowly zooms in on her from above. 0:03-0:07 closer shot of her face as she turns her head. Simple cinematic zoom in, whatever framing looks best. ACTION: She breathes deeply and turns her head. Her hands move expressively on the carpet, fingers curling, flexing and grasping the fabric, emotional hand gestures throughout. ANATOMY AND CONTACT: Delicate elegant hands, beautiful fingers, gold bracelets shimmering and glinting as she moves. Her long hair spreads and flows across the carpet. LIGHTING: Beautiful moody cinematic lighting, dramatic shadows, soft glow, atmospheric dark room with red tones. AUDIO: Heavy breathing, cinematic ambient. LOOK: Cinematic, moody, dramatic, film noir vibe, award winning, 8k, hyperrealistic.

In the red carpet scene, "moody" and "dramatic" were the only direction the face got, and what came back is a woman lying still with nothing readable on her.

Good prompt: the same mood carried by a visible breath, a gaze, and still hands.

SCENE: A young woman lies on her back on a deep red carpet in a dim room. She draws one slow deep breath and turns her head. One continuous take, about 7 seconds, real-time. SUBJECT: Woman 27, shoulder-length waves fanned on the carpet, ribbed cream long-sleeve top, patterned silk scarf at the neck, large gold hoop earrings, one heavy gold bracelet on the right wrist. Face, scarf pattern and jewellery stay identical for the whole take. CAMERA: Single continuous arc push-in, no cuts. Starts at a medium high angle showing her torso and outstretched arms, then descends and swings counter-clockwise around her head into an extreme close-up on her eyes, finishing at a lower side angle. Focus rides the eyes throughout, foreground hand allowed to fall soft. 40 degree field of view, shutter 270-360 degrees, slow constant speed, no zoom snap. ACTION: Her chest rises on a long inhale, lips parting slightly, then she turns her head to camera-right and lifts her gaze past the lens. Both hands stay still and relaxed on the carpet, palms up, fingers loose and unmoving. No twitching, no curling, no gesture. ANATOMY AND CONTACT: Each hand keeps five separate anatomically correct fingers with natural joint angles and an unbroken wrist line; no bending past the natural range, no fusing with the carpet. The gold bracelet stays a rigid ring locked to the wrist, sliding only as far as gravity allows, its engraving fixed. Hair rests on top of the carpet pile with real contact and slight strand displacement as the head turns, never passing through the fibres. LIGHTING: Low key. Soft warm key from above and behind the crown of her head, modelling the cheekbones and leaving the eye sockets in shadow, almost no fill, high contrast. One deep red practical glowing in the background left. Light direction and intensity stay constant as the camera arcs. No flicker, no added flare. AUDIO: One heavy slow inhale and exhale, synced exactly to the chest rise and the parting lips, over a low room drone. No music, no dialogue. LOOK: Photoreal, shot on ARRI Alexa 35, split-toned grade of velvet red, cold swampy green shadows and warm gold speculars, fine grain, visible carpet pile, knit rib and skin micro-texture. Background geometry stays rigid through the parallax. No text, no logos, no watermarks.

Name the visible action instead of the feeling: a held breath, a gaze lifting past the camera, a mouth that doesn't quite smile. Camera movement carries mood too, worth describing with the same care as the face. A handheld shot with a slight wobble often reads as more human than a perfectly smooth glide, since a little imperfection in the camera itself can sell a mood that a static, polished frame won't.

Why does lighting and atmosphere miss the mood I wanted?

"Moody" and "cinematic" leave the physics of the scene open. The model still needs a source, a time of day, and an angle, and it picks its own if the prompt doesn't, which is how a scene meant to feel tense or intimate comes out flat or oddly lit.

Bad prompt: a rally car with the light and the time of day never named.

Epic rally car racing through a foggy forest on a dirt road, super cinematic, insane, 4k, hyperrealistic, best quality. 0:00-0:02 wide shot of the car driving fast. 0:02-0:04 car hits a puddle, huge massive mud explosion everywhere. 0:04-0:06 inside the car, driver in a helmet, wipers going. 0:06-0:08 side shot of the car, dynamic. 0:08-0:13 the car flies off a big jump, hanging in the air, epic slow motion moment. 0:13-0:15 camera on the hood driving forward. Dynamic cinematic angles, choose the best angle for each shot, make it intense. Loud engine roar the whole time. Moody teal grading, atmospheric fog, epic mood.

Good prompt: overcast dusk as the key, headlights as the only practical, screen direction locked.

SCENE CONTEXT: Rally special stage on a wet gravel forest road, Nordic pine woods, low fog, overcast dusk. One rally car runs the stage. Real-time, about 15s, seven segments, six hard cuts. SUBJECT LOCK: One rally car only, white and blue livery, round door number panel, mud-caked flanks, headlights on. Right-hand drive, steering wheel on the right in every exterior and interior shot. Driver in helmet with closed dark visor and gloves, co-driver right of frame in the cockpit. Livery pattern, mud level and light state stay identical across all cuts. SCREEN DIRECTION LOCK: The car travels right-to-left in every exterior shot. Camera stays on one side of the road; the 180 axis never flips. Fog density and road surface stay continuous across cuts. BEATS: 1. Wide static, 47 degree FOV, car exits a bend right-to-left, headlight cone on wet gravel. Hard cut. 2. Wide, tighter framing and lower angle, car hits a standing puddle, fine dispersed water and mud droplets burst outward and fall fast, thin spray that clears within half a second, silhouette of the car never hidden by a solid wall of mud. Hard cut. 3. Cockpit medium, camera on the roll cage, driver right of frame at the right-hand wheel, wipers sweeping in rigid synchronous mechanical arcs across muddy glass, rain beads streaking. Hard cut. 4. 29 degree FOV side tracking pan, car right-to-left, background streaked with motion blur, car sharp. Hard cut. 5. Extreme close-up, 84 degree FOV fender-mounted, tread biting wet gravel, stones and mud flicking off the tyre. Hard cut. 6. Wide static, car launches over the crest right-to-left, a fast clean parabola, apex reached in under half a second, immediate descent, heavy landing with suspension compressing hard, gravel and spray kicked out on impact, car continues out of frame. No hovering, no mid-air pause, no slow motion. Hard cut. 7. Bonnet POV, 84 degree FOV, road rushing forward through fog, final static held beat as the road straightens. PHYSICS: Full gravity and inertia at all times, mass loads the outside suspension in the bend, the body rolls and settles, tyres deform under load. Water breaks into droplets, never into a frozen solid cloud. Mud lands and sticks, spray thins as it rises. Wipers are rigid mechanical arms with a fixed pivot, no bending, no morphing. CAMERA AND OPTICS: Static locked-off or single-axis pan only, camera height 60-90 cm. Shutter 270-360 degrees, long natural motion blur. No zoom, no whip pans, no digital jitter, no effects. LIGHTING: Overcast 6000K daylight through fog as key, soft and directionless. Headlights as the only warm practical, cutting a visible cone in the mist. No unmotivated sources. AUDIO: Turbocharged four-cylinder engine, real gear changes and anti-lag pops. Distinct impact sounds, puddle slap on beat 2, engine rising as the wheels leave the ground on beat 6, hard suspension slam and gravel scatter on landing. Engine level drops with distance and loses high end when the car is far or out of frame. No music. LOOK: Photoreal, shot on ARRI Alexa 35, natural grain, cool teal-green forest grade, wet gravel and mud texture, condensation on glass. No text, no logos beyond the car livery, no watermarks.

Name the light source, the angle, and the time of day directly. Headlights cutting through rain at dusk carries more atmosphere than the word "atmospheric" ever will. In Cinema Studio this stops being a sentence at all: genre, lens, and lighting get set once as parameters and hold across every shot.

Why does movement bend or warp in AI video?

A single adjective leaves the entire path of a move undefined. "Flies fast" says nothing about where the turn starts, where it ends, or how sharp it is, and that gap is exactly where a rigid aircraft ends up flexing mid-turn like it's made of rubber, or a clean shot picks up a sharp, unplanned zoom at the very end that undoes everything that worked before it.

Bad prompt: a fighter jet with the flight path described as an adjective.

SCENE: A futuristic fighter jet flying through a huge sci-fi megacity. Epic, cinematic, about 10 seconds, insane visuals. SUBJECT: A cool white and red space fighter with glowing blue engines and a pilot inside. Super detailed, 8k, hyperrealistic, best quality. CAMERA: 0:00-0:02 the jet flies past the camera. 0:02-0:05 chase shot behind the jet, epic sun in the lens. 0:05-0:10 the jet banks hard and flies away. Dynamic cinematic angles, cut between the best angles, make it intense. FLIGHT PATH: Show the jet flying in different directions for maximum dynamism, weaving between skyscrapers, whatever looks coolest each shot. PHYSICS: The jet bends and flexes as it banks, fluid dynamic motion, the whole ship moving organically through the turn. Glowing engines, bright blue thrusters. LIGHTING: Gorgeous sunlight bursting into the lens with epic anamorphic flares in every shot, dramatic light changes, moody and cinematic. AUDIO: Loud epic jet engine roar the whole time, cinematic sound design. LOOK: Cinematic, epic, sci-fi, award winning, 8k, hyperrealistic, insane detail.

The jet prompt asked for the hull to "bend and flex" through the turn and got exactly that.

Good prompt: one continuous path, a rigid hull, a single sun angle.

SCENE: A single-seat futuristic fighter runs low through the canyon between megacity towers. One continuous take, about 10 seconds, real-time. SUBJECT: One fighter only, white-grey worn hull with red stripes on nose and wings, twin engine nozzles glowing blue, glazed canopy with a helmeted pilot visible. Hull is rigid metal, panel lines, scratches and proportions stay identical from first to last frame. CAMERA: Single continuous chase, no cuts. Starts rear-three-quarter left, arcs around the fuselage as the jet banks, passes under a skybridge, settles dead astern by the end. Camera speed constant, subtle air-turbulence micro-shake only. 47 degree field of view, shutter 270-360 degrees, no zoom punches, no speed ramps. FLIGHT PATH: The jet travels screen-left to screen-right for the whole take and never reverses direction. It banks left through a wide arc, levels out, drops under the skybridge, then accelerates straight into depth. Towers, bridges and traffic lanes pass at a consistent rate that reads the speed. PHYSICS: Rigid-body aircraft, nose, wings, tail and nozzles move as one solid unit through the bank, zero flex, zero independent drift of the tail section. Wingtips hold exact length and dihedral at maximum roll. Engine exhaust pushes a visible heat shimmer that distorts the towers behind the nozzles, with a faint ionized wake. Background traffic ships keep continuous trajectories, none pop in or vanish. LIGHTING: One motivated key, hard sun high and camera-right, held at the same angle for the entire take, raking the upper hull and leaving the lower fuselage in cool bounce from the glass towers. Passing under the skybridge, the jet and the frame darken and recover naturally. Anamorphic blue flare appears only when the sun is actually in frame, never as a random burst. AUDIO: Low turbine roar with the source panned to follow the jet across the stereo field, dopplering as it swings. Under the bridge the sound goes muffled and reflective for a moment, then opens up. Distant city hum underneath. No music. LOOK: Photoreal, shot on ARRI Alexa 35, cool grey-blue metropolis grade against warm sun, soft contrast with detail held in shadow, fine grain, real metal wear. No text, no logos, no watermarks.

Describe the turn itself: start point, end point, how sharp. Cinema Studio's motion settings hold that path fixed instead of re-guessing it on every run, the same way a real camera rig would. Paired with a model built for motion, like Seedance, a dynamic sequence holds together more reliably end to end.

Why does my AI video feel cluttered and lose its focal point?

A shot list built from angle names alone doesn't say what each shot is actually for. Wide shot, dynamic angle, close-up, none of that tells the model what to prioritize, so frames pile up that don't serve the scene, an odd fisheye angle, a shot that lingers on a person's silhouette when a product was supposed to be the point of the whole video.

Bad prompt: eleven shots with no job assigned to any of them.

SCENE: Epic mountain bike commercial, studio shots then city then forest then mountains. Eleven shots, about 15 seconds, insanely cinematic. SUBJECT AND IDENTITY: A beautiful green mountain bike with black parts and a cool rider in casual clothes. Premium product, 8k, hyperrealistic, best quality. SHOTS: 0:00-0:02 macro shots of the chain and the tire. 0:02-0:03 the bike in a clean studio. 0:03-0:05 rider riding fast in the city, then in the forest. 0:05-0:09 he rides down the trail, wheel hits a root, epic action. 0:09-0:11 close ups of the brake and the gears. 0:11-0:15 he rides a mountain road and stands on the summit at sunset. Whatever angle looks coolest for each shot. ANATOMY AND BIOMECHANICS: Close up of a hand squeezing the brake lever, fingers wrapping tightly around the metal. Legs pedaling furiously, blurred fast motion of the legs, dynamic energetic cycling. PHYSICS: Wheels smashing through roots and rocks, huge explosions of dirt and dust flying everywhere, epic particles, dramatic energy in every shot. CONTINUITY: Show the bike from lots of different angles and directions to keep it dynamic, jump between locations for variety. LIGHTING: Beautiful cinematic lighting in every shot, gorgeous sun flares, moody atmosphere, epic golden hour vibes. AUDIO: Energetic music, bike sounds, epic sound design. LOOK: Cinematic, premium, epic commercial, award winning, 8k, hyperrealistic, insane detail.

The bike commercial asked for "whatever angle looks coolest" and ended on the rider instead of the bike it was selling.

Good prompt: eleven numbered shots, one job each, screen direction held throughout.

SCENE: Commercial spot for one mountain bike, studio detail shots, then city, forest trail, mountain switchback and a final hero shot. Eleven segments, ten hard cuts, about 15 seconds, real-time. SUBJECT AND IDENTITY: One bike only, olive-green frame, black components, matte finish, brand engraving on the top tube, knobby tyres, disc brakes. Frame geometry, shock mount, cable routing and decal placement stay pixel-identical in the studio shots and in every riding shot. Rider, man 30, white tee, olive overshirt, dark trousers, white trainers, same clothing throughout. SHOTS: 1. Macro pan across chain and chainring, teeth engaging link by link. 2. Macro dolly along the tyre tread. 3. Studio side tracking of the whole bike on a plain light grey cyclorama, bars turned slightly camera-right. 4. Low wide, rider riding screen-right to screen-left on city asphalt against the sun. 5. Chase from behind and above as he drops into a forest trail. 6. Low close-up of the rear wheel rolling over a tree root. 7. Medium tracking of the descent, still screen-right to screen-left. 8. Detail of the hand squeezing the brake lever. 9. Detail of the derailleur shifting across the cassette. 10. Rear three-quarter as he leans into a switchback. 11. Wide, rider standing beside the bike on the summit, back to camera, long held final beat. ANATOMY AND BIOMECHANICS: On the brake shot the hand keeps four separate fingers wrapped around the lever with visible knuckles and nail edges, the index and middle finger pulling while the thumb stays hooked under the grip; no finger sinks into or fuses with the metal. Pedalling is biomechanically correct, soles stay fixed on the pedals, ankles articulate, cadence matches the wheel speed, legs never intersect the frame or slide free of the cranks. Through the switchback the outside pedal is down and loaded, the inside pedal up, feet still. The rider's body works constantly, knees and elbows absorb impacts, hips shift back on the drop, shoulders steer, torso moves relative to the bike. PHYSICS: Tyres deform against roots and rocks and roll over them with real contact; nothing passes through solid geometry. Dirt, gravel and dust are three-dimensional particles with weight, thrown off along the tangent of wheel rotation, arcing and falling under gravity, settling on the ground. Wheel axles rotate on a fixed centre with no jitter or position jumps. Suspension compresses and rebounds visibly on every impact. CONTINUITY: The rider travels screen-right to screen-left in all riding shots. The location changes only where the terrain plausibly connects, city to forest edge to mountain road. Asphalt tone, texture and grade match across the cut between the wheel detail and the wide switchback. Background geometry, foliage, rocks and roadside stay rigid and stable during fast pans, no melting, no shifting textures. LIGHTING: Studio shots, soft even diffused light, clean grey background, controlled contrast. City, overcast daylight, low sun behind the rider. Forest, warm side and back light broken by the canopy. Mountain, hard midday sun with defined shadows. Final shot, low golden sun behind the rider. Each location keeps one consistent motivated direction. CAMERA: Macro shots at 29 degree field of view with shallow focus; riding shots 47 to 84 degree, camera 0.5 to 3 m; smooth dolly, crane and vehicle-tracking moves, shutter 270-360 degrees, natural motion blur, no speed ramps. AUDIO: Freehub ratchet, tyre hum on asphalt, gravel crunch and root impact, brake lever click, crisp derailleur shift, wind on the descent. Each effect lands exactly on its visible action, over a driving music bed cut to the edit. LOOK: Photoreal, shot on ARRI Alexa 35, natural anamorphic character, deep greens and earth tones, controlled saturation, fine grain, matte paint, machined metal, real rubber. No text, no logos beyond the bike engraving, no watermarks.

Give every shot one job, this one shows the frame, this one shows the rider, this one shows the product in hand, and drop anything that doesn't serve it. Go back through the result frame by frame and cut what doesn't earn its place. If the product needs a moment fully to itself, Marketing Studio can generate a few product-only passes separately, which then drop into the final cut through Canvas.

Why do characters lose their positions in complex AI scenes?

The overall idea of a scene isn't the same as knowing where everyone stands. A touchdown catch, a group around a fire, describing the idea leaves position, background action, and camera timing wide open, and a sequence that should feel dynamic and controlled turns into drifting positions and characters who seem to forget where they were standing a moment ago.

Bad prompt: a touchdown described as an idea, with no positions for anyone.

SCENE: Epic American football game at night under stadium lights, huge touchdown moment. Ten shots, about 16 seconds, insanely cinematic. CAST AND IDENTITY: Lots of players in white and red uniforms, a quarterback, a receiver, defenders. Each shot shows a different hero player up close. Beautiful athletes, 8k, hyperrealistic, best quality. SHOTS: 0:00-0:02 wide shot of the teams on the line. 0:02-0:04 close up macro of hands gripping the grass, fingers digging deep into the turf. 0:04-0:06 the lines crash, players smashing and crashing through each other. 0:06-0:10 quarterback throws, receiver runs. 0:10-0:13 the ball flies straight into his hands in the end zone. 0:13-0:16 epic aerial shot of the whole stadium going wild. PHYSICS: Massive brutal collisions, bodies flying, players sliding across the field in dramatic slow motion as they fall. Explosive energetic movement everywhere. CROWD AND SIGNAGE: Packed stadium, fans cheering with raised hands in the foreground, team name written big in the end zone, banners and scoreboard text everywhere, sponsor boards around the field. LIGHTING: Gorgeous stadium lights with epic anamorphic flares in every shot, dramatic god rays, moody atmospheric haze, cinematic light. AUDIO: Epic orchestral music, crowd going crazy, big impacts. LOOK: Cinematic, epic sports drama, award winning, 8k, hyperrealistic, insane detail.

The football prompt described the touchdown as an idea rather than a set of positions, and players slid through each other on the way to it.

Good prompt: six numbered shots, named players, solid-body collisions, identical lighting across cuts.

SCENE: Night American football game under stadium floodlights. A snap, a line collision, a long pass and a leaping touchdown catch. Six segments, five hard cuts, about 16 seconds, real-time. CAST AND IDENTITY: Two teams only, white-and-navy offense, deep red defense. The hero receiver is number 44 in white, red gloves, scuffed silver helmet, mud on the left thigh. The quarterback is number 12 in white, same face, chin strap and glove colour in every shot he appears in. Numbers, helmet decals and mud pattern stay identical across all cuts; no player changes number or build mid-action. SHOTS: 1. Wide static, grass-level, both lines set in three-point stance, breath fogging, crowd behind. Hard cut. 2. Macro of the snapper's right hand planted on the turf, five separate anatomically correct fingers with natural joint angles and normal thickness, fingertips resting on the grass and lifting cleanly as the ball snaps. Hard cut. 3. Tracking pull-back at knee height as the lines collide, camera weaving between running legs. Hard cut. 4. Tilt up on the ball spiralling against the floodlight masts, clean stable rotation axis. Hard cut. 5. Medium wide, number 44 leaps, takes the ball, is hit by a defender and rolls onto the turf. Hard cut. 6. Low-angle medium from behind number 44 standing, fists raised, haze and rim light, long held final beat. PHYSICS: Bodies are solid and heavy, shoulder pads meet and stop each other, no player passes through another, no mesh overlap on contact. Every collision transfers momentum, staggers the smaller man and compresses the pads. Cleats bite the turf with real grip, tearing divots, no sliding feet. The ball decelerates into the hands, the gloves close around it and absorb it with a visible squeeze, then it stays locked to the forearm. The fall carries full body weight, impact, bounce of flesh and pads, grass and dirt kicked up, limbs settling from inertia. CROWD AND SIGNAGE: Stands filled with individually moving spectators, varied posture and timing, camera flashes at random intervals. Any foreground hands have correct separate fingers. Field markings are only plain white yard lines and hash marks. No lettering, no team name in the end zone, no banners, no scoreboard text anywhere in frame. LIGHTING: Four floodlight masts high and behind the action as a constant hard backlight, rimming helmets and breath, faces in shadow under the facemasks. Light haze makes the beams volumetric. Direction and intensity stay identical across every cut. Anamorphic blue flare only when a mast is actually in frame. CAMERA: Grass to knee height throughout, 47 degree field of view on wides, 84 degree on the macro and the leg-level tracking, shutter 270-360 degrees, natural motion blur, operator-held weight, no speed ramps, no slow motion. AUDIO: Quarterback cadence call, pad and helmet impacts with real low-end, cleats tearing turf, whistle on the catch, crowd roar rising into a wide reverberant stadium bowl. Score builds under it. Impact sounds match the visible force. LOOK: Photoreal, shot on ARRI Alexa 35, cool desaturated green turf against warm floodlight, lifted matte shadows, 35mm grain, visible grass blades, leather ball grain, helmet scratches. No watermarks, no logos.

Break the scene into individual shots and write out each person's position and the background action for every single one, not just what the main character is doing. For anything this layered, Popcorn plans the whole sequence as a storyboard first, positions and lighting locked in, before any of it gets animated into motion.


How do I improve motion and lighting beyond the prompt?

A prompt describes motion, and some of it comes across more clearly as a setting. Two tools inside Higgsfield's AI-native creative suite work alongside the prompt for exactly that. Cinema Studio is how Higgsfield handles cinematic video production: camera, light, and motion are chosen before generation starts, so a movement carries a set speed and path of its own, with no need for the model to read that off words like "fast" or "dynamic."

  • Genre sets the physical register first, seven options from Drama to Noir

  • Lighting works as a fixed preset instead of a described mood, seven presets, Contre-jour for a rim-lit silhouette among them

  • Camera MoveSet Style turns motion into a real setting across ten options

  • Lens, focal length, and aperture round it out: six lens characters, five focal lengths, three aperture values

Popcorn is how Higgsfield handles storyboards and multi-frame sequences: a set of frames that needs to hold together as one.

  • Generates up to eight frames as one coherent set

  • Manual mode directs frame by frame, Auto mode expands one prompt into the full sequence

  • Draws on up to four references, so continuity comes from shared material rather than being rebuilt from text each time

  • The frames are the plan, not the finished video, a model animates them afterward, chaining the last frame into the next to keep motion continuous


Checklist Before You Generate

The steps above are how a prompt gets built. This is the pass to run over it before hitting generate.

  • Character detail: age, build, and clothing, backed by a reference or asset

  • Points of contact: hands, fingers, and anything touching something else named explicitly

  • Emotion: tied to a visible action, not a mood word

  • Light: source, angle, and time of day named, or set in Cinema Studio

  • Movement: an explicit turn, tilt, and angle, not an adjective

  • Shot purpose: one job per shot, nothing in frame that doesn't serve it

  • Complex scenes: storyboarded in Popcorn first

We walk through several of these on camera, with the generations side by side, in AI Videos Look Bad? Here's Why.

What do all these mistakes have in common?

Plastic characters, fused hands, flat emotion, warped movement, camera clutter, drifting positions: seven different symptoms, one cause. Something in the prompt was left open, and the model closed it on its own.

The habits that fix all seven are the same few. Name each person and back it with a reference. Say where objects sit and how hands make contact. Tie emotion to something visible. Name one light source with a direction. Write movement as a start point and an end point. Give each shot one job, and block the positions before generating a complex scene.

Cinema Studio and Popcorn hold the parts a sentence carries least easily, the light's position, the camera's path, a sequence's continuity, set once and kept across every shot. What stays in the prompt is what a prompt does well: who's in the scene, what they do, and how it should feel.

Common AI Video Prompt Mistakes (And the Fixes That Actually Work)

Try Cinema Studio

Got any questions left?

Yes, just for different things. Cinema Studio on Higgsfield holds camera, lighting, and motion as settings, which leaves the prompt to carry who's in the scene, what they do, and how it should feel.

A prompt problem changes on every run. A settings problem is wrong the same way every time, and that one is usually quicker to fix in Cinema Studio than to write around with more text.

Not a short clip with one clear action. Popcorn earns its place once a scene has several characters or several beats, and its frames feed straight into the video models on Higgsfield.

Train the face once. Soul ID on Higgsfield builds a persistent identity from a set of reference photos and applies it to later generations, so the character holds without being described again each time.

Often. Most of these come from one thing, a detail left for the model to invent. Setting the lighting in Cinema Studio, for instance, tends to clean up the mood and atmosphere in the same pass.

by Higgsfield

Share article

Discover more

View all