Help center

How do I write a good prompt?

Updated

A prompt is how you tell the model what to create, and the single biggest thing you can do to get a better result is simple: be specific. A short, vague prompt leaves the model to guess, and it will guess differently every time. A detailed prompt that spells out the subject, the details, the setting, and the style gives you something much closer to what you pictured. How you prompt also depends on where you start: describing a scene from words works differently from animating an image you already have. This guide covers both. And if you'd rather not write a detailed prompt yourself, Supercomputer can take a rough idea and finish the job for you.

Can Supercomputer just do it for me?

Yes. If you have a loose idea and don't want to polish a prompt, describe it in plain language in Supercomputer, even something as vague as "make a TikTok for my sneakers." It fills in the details, checks the plan with you, and generates the result, routing each step to a suitable model automatically. It's the fastest path from idea to finished asset. See How do I use Supercomputer?

For everything else, writing your own prompt gives you more control. Start here.

Why does being specific matter so much?

The difference between a weak result and a strong one is usually detail. Think of a detailed prompt in layers: subject (who or what is in the frame), details (appearance, clothing, expression), environment (where it happens), and style (photorealistic, cinematic, and so on). For video, add camera and mood.

You don't have to fill every layer every time, but the more precisely you describe what matters, the more control you have. The best way to see this is to follow one character through a full workflow.

Step 1: Build a base character

Start with a clean, detailed description of who your character is. A neutral studio shot with a plain background gives you a clear base you can restyle later.

Photorealistic full-body studio portrait of a rugged man in his early 50s, standing straight and facing the camera, arms relaxed at his sides. Short cropped graying hair, weathered tanned face with deep wrinkles, gray stubble beard, intense stern gaze. He wears a plain heather-gray t-shirt, dark gray cargo pants and black sneakers. Clean seamless white studio background, soft even lighting, sharp detail, visible skin pores, shot on medium format camera. Full body visible head to toe, centered composition.

Notice how much is specified: age, hair, skin texture, exact clothing, background, lighting, even the camera type. That precision is what makes the character reusable.

Step 2: Restyle the same character

To change the look while keeping the same person, tell the model explicitly what to keep and what to change. Lock the face, pose, and background, then change only the wardrobe.

Keep the exact same man, same face, same pose, same white studio background and lighting. Change only his outfit: dress him as a Wild West gunslinger, black wide-brimmed cowboy hat, long worn black leather duster coat, dark blue button-up shirt, black leather bandolier with bullets across his chest, black leather gloves, dark leather chaps, black cowboy boots. He holds two silver revolvers crossed at his waist. Stern menacing expression.

The phrase "keep the exact same man, same face, same pose" is doing the heavy lifting: it's what holds identity steady while everything else changes.

Step 3: Place the character in a scene

Now move the character into a real location. Describe the setting in cinematic terms, lighting, depth, mood, so the character feels part of the world rather than pasted onto it.

Place this exact cowboy character inside a dim 19th-century Western saloon. Cinematic close-up of his face under the black cowboy hat: weathered wrinkled skin, gray stubble, dirt and dust on his cheeks, cold intense stare slightly off-camera. Soft warm light from a window behind him, blurred saloon interior in the background, wooden beams, hanging lamps with warm glowing bokeh, hazy dusty air. Moody cinematic color grading, shallow depth of field, film grain, anamorphic look.

This frame becomes the starting point for your video: the character and the location are now locked in a single image.

Step 4: Bring the scene to life with video

Here's where prompting changes completely. When you animate an image, the picture already holds the character, the setting, the lighting, and the style, so your prompt stops describing what things look like and starts describing what happens: action, camera movement, timing, and mood.

A video prompt is more like a shot list than a sentence. For a cinematic clip, describe it in blocks:

  • Style and mood — the overall tone (here: dry neo-western comedy, deadpan calm against quiet fear)

  • Cinematography and camera — lens choices, movement, framing

  • Lighting and color — the palette and how the scene is lit

  • Action — beat by beat, what the character and the scene do

  • Audio — score, ambient sound, and any dialogue with timing

  • Shot list — the sequence of shots with durations

For example, the cowboy clip is built as 6 shots over 15 seconds: the re-entrance, a flat threat delivered to the room ("Somebody took my horse... you've got twenty minutes"), the patrons' uneasy reactions, and the cowboy calmly pouring a drink. Each shot specifies its lens (35mm for the room, 75mm for faces), duration, and action.

Style: Live-action photographic realism, dry neo-western comedy, the cowboy walks back in and flatly gives the bar twenty minutes to return his horse, with a vague deadly threat, then calmly pours another whiskey while the room quietly panics. Warm, dusty, cinema-grade, deadpan against real fear. NOT CGI-looking, NOT plastic, NOT a commercial. References (in spirit, fully original): the quiet menace of a calm man in a classic Western thriller. Calm, threat, restrained fear. Cinematography: A continuous scene, the re-entrance, the flat threat delivered to the room, the patrons' uneasy reactions, the cowboy sitting and pouring again; composed coverage from real positions, moving forward. NO body-mounted rig. LENS DISCIPLINE: spherical rectilinear, 35mm for the room, 75mm for faces, NO fisheye, NO warp. Eye-line: the cowboy addresses the room flatly; the patrons dare not meet his eye, never the lens. Lighting: Naturalistic, motivated, warm and dusty, golden window light, hazy air, deep shadow, faint neon. Carves the cowboy's calm face and the patrons' uneasy ones, keeps the dread close. Deep warm contrast. NOT bright, NOT flat, NOT studio. Color: Warm, dusty, muted, filmic, 60% amber and ochre, 30% deep shadow, 10% accents (faint neon glow). Filmic print grade, fine grain, gentle halation. NOT teal-orange, NOT saturated, NOT digital-clean. Camera: Cinema-grade digital camera with spherical prime lenses at 35mm and 75mm, T2.0 to 2.8, dolly, Steadicam, or sticks, never body-mounted. Naturalistic depth, rectilinear drawing. 24fps, real-time throughout. Fine 35mm-style grain, constant. Skin: Anti-plastic, pore-level realism, real weathered matte skin, NOT glossy. The cowboy: the same older deadpan man; the patrons: weathered working men and a bartender, now uneasy and frightened but restrained; all ordinary human eyes, NO glowing eyes, NO eye-shine. NOT smoothed, NOT waxy. Acting: Deadpan calm against quiet fear, the cowboy walks back in, stops, and flatly tells the room someone took his horse and they have twenty minutes to return it, or he'll do what he did in Texas ten years ago; the room goes still, patrons exchanging uneasy frightened looks, a swallow, men freezing, genuine fear played restrained, not slapstick; then the cowboy calmly sits, pours another whiskey, and drinks, utterly unbothered. The contrast of his calm and their fear is the comedy. Natural micro-life: his flat delivery and steady sip; their nervous stillness, a bead of sweat, darted glances. Nobody mugs at the lens. Physics: Honest, the door swings, boots on wood, the room stilling, a chair creak, whiskey poured, a glass set down. Nothing floats or loops. Composition: Full-frame 16:9, fills the entire frame edge-to-edge, NO letterbox, NO black bars, NO matting, NO cropped strips. The opening holds his re-entrance and the room; the scene moves between his flat delivery, the uneasy faces, and his calm pour. Layered depth: a foreground table or glass, the cowboy mid, the frightened room behind. Continuity: One bar, one cowboy, one continuous threat-and-pour, wardrobe, the bar and warm light consistent. Clean hard cuts only, NO morphing transitions, NO position jumps. Editing: Deadpan-vs-fear rhythm, the re-entrance, the flat threat, the uneasy room, the calm pour. 6 shots in 15 seconds, cut to a sparse tense-wry score. Clean hard cuts only. NO transitions, NO fades, NO speed-ramps. Technical: Full-bleed 16:9, no letterboxing, bars, or cropping. 24fps, real-time. Rectilinear 35mm and 75mm lenses, fine grain, warm dusty naturalistic filmic grade. NO body-mounted camera. The threat is the cowboy's single-speaker delivery with lip-sync priority (mouth lit and in focus, slow, clear, spaced); patrons react non-verbally. NO on-screen text, NO subtitles. NO glowing eyes, NO fisheye, NO plastic skin, NO commercial gloss. Non-graphic, no violence, the threat stays a threat. Strictly non-IP: neon and signs indistinct, no brand logos, gear generic, no readable text, no real-location ID, every surface clean or an indistinct unreadable mark, zero AI-text artifacts. Audio: Original instrumental score, fully generic, NOT mimicking any artist or track, NOT quoting any melody. A sparse, tense-but-wry neo-western cue, a low drone and a lonely guitar, near-silence holding the threat. Diegetic bed: the door, boots, the room going dead quiet, a chair creak, whiskey pouring, a glass set down. On-camera dialogue, English, the cowboy only, flat and unhurried, single-speaker (lip-sync priority): COWBOY, around 2.4 to 4.0 seconds: "Somebody took my horse." COWBOY, around 5.0 to 7.4 seconds: "You've got twenty minutes to bring it back." COWBOY, around 8.0 to 10.6 seconds: "Or I'll do what I did in Texas. Ten years ago." No other dialogue, patrons react non-verbally. No copyrighted music. No recognizable melody. No real-brand audio. Character description: The cowboy, the same older weathered deadpan man, calm as stone, issuing a vague deadly threat; the patrons, weathered working men and a bartender, now quietly frightened but restrained. Grounded; the comedy is the contrast. Location: The same dim, warm, dusty dive bar, cluttered indistinct walls, faint neon, golden hazy light, deep shadow, a worn bar, mismatched tables, a pool table. Lived-in, cinematic. Strictly generic, no readable signage, no IP, no real-location ID. Hero prop: The threat itself, and the whiskey he calmly pours right after it; the frightened still faces. No readable text anywhere, neon and signs indistinct, no brand logos, generic glassware, no real-location markings; every surface clean or an indistinct unreadable mark, zero AI-text artifacts. Mood and tempo: A calm man threatens a whole bar and then pours a drink, 15 seconds, full-bleed 16:9, 6 real-time shots, warm dusty light, a flat threat and a frightened still room, ending on his unbothered sip. Dry, tense, funny, cinematic.

The more precisely you direct each beat, the closer the result. You don't need every block for a simple clip, a single line of motion works too, but for cinematic sequences, structure is what gives you control.

How do I keep the same character across generations?

The workflow above works because it's one consistent character, and that isn't automatic. Generate the character once, then reuse it as a reference so the face and design stay the same. For a real person, this is Soul ID; for any character carried across models, save it as an Element and reference it in later prompts. See How do I create and use a Soul ID character?

What habits make any prompt better?

Say what you want, not what you don't. Positive descriptions generally work better than long lists of "no this, no that," though naming a few key things to avoid (like "no readable text" or "no plastic skin") can help on detailed shots.

Keep it focused for simple tasks. A short clip doesn't need a six-block shot list. Match the prompt's complexity to the result you want.

Avoid contradictions. Make sure your prompt and your input image agree: a "still, calm" instruction fights a motion-blurred source frame.

Write in English. Models are trained mostly on English, so it gives the most predictable results. One exception: Chinese-developed models like Seedance and Kling sometimes respond better to Chinese for nuanced motion.

How do I improve a result I don't like?

Treat prompting as a conversation: generate, look, change one thing, generate again. Changing a single element at a time, the lighting, the lens, one detail of wardrobe, teaches you what each change does. To build a longer sequence, use the last frame of one clip as the starting image for the next, then join them in an editor.

Need inspiration?

Browse generations in the Higgsfield Community and study the prompts behind results close to what you're trying to make.

Share Feedback

Was this article helpful?

by Higgsfield

Share article