Scenes 1-2: The Opening & Transformation

With every asset locked and the prompting framework primed, it's time to shoot the first two scenes of the ad — the moment the story actually starts.

Cover still from the opening scene: the hero reaches for the floating can on the empty New York street

Scene 1 — the floating can

The opening is deceptively simple: a guy walks down an empty city street, notices a can floating in front of him, and it lands in his hand. The prompt is kept short by design, but the scene is treated as the most important one in the whole ad — if the first shot feels off, viewers stop watching.

The prompt is written short on purpose: the walk, the can appearing, it landing in his hand, and a note to keep the camera handheld with a shaking effect. Everything else — the lighting, lens, acting style, and physics — carries over automatically from the style header set up in the prompting framework, so every future shot stays visually consistent without having to redescribe it each time. Also baked into the header: no music, environmental sound only, since the score gets added later in the edit.

Four batches come back, and the first one already has it all — handheld camera-shake that reads as shot-by-hand rather than generated, real sun flares hitting the lens, and a slow-motion moment as the can drops into his hand. There's also a bonus: the earlier test clip from Stage 1 — the one used to check that the character and location worked together — turns out good enough to reuse as the opening beat itself. The final Scene 1 is cut together from the best moments of both.

Scene 2 — the transformation

This is the big one: a close-up of the hero drinking the soda, a big smile, then he sprints and his robot armor builds up around him mid-air. If this shot doesn't feel epic, the story doesn't land.

The first pass, run in four batches, misses on nearly every count: a broken smile, the character running backwards instead of forward, in another take his eyes go blank and shift color out of nowhere, and in yet another the transformation itself falls flat — he just stands there instead of moving through it. None of it is usable — but instead of manually rewriting the prompt, the fix happens by talking to Claude like a director giving notes: keep moving fast through the transformation, add creative handheld camera movement and angles, make sure he runs forward, and fix the face — natural smile, alive eyes, no color change.

That single note produces a version that finally works: the camera feels alive, the transformation reads as continuous motion, and the face holds up. One more round of notes pushes it further, asking for extreme close-up "impact" shots — macro shots on the armor pieces locking into place: the chest plate clamping, the wrists snapping in, the helmet sealing shut. That pass nails it, and the final Scene 2 is assembled from the best moments across all the generations.

The pattern across both scenes: prompts don't get hand-edited line by line. You watch the result, describe what's wrong in plain language, and let Claude rewrite the prompt around your notes.