Sora 2 Prompting Guide
Sora 2 is a breakthrough in AI video generation. Here's our guide to getting the best out of your prompts.
Sora 2 raised the bar on AI video generation, not just on visual quality but on how well it follows detailed, specific direction. Getting good results out of it means prompting more like a film director and less like you're describing a picture. Here's how to do that.
Section 01: Think in Shots, Not Just Scenes
The single biggest upgrade to your Sora prompts is describing a shot the way a cinematographer would: subject, action, camera movement, and setting, in that order. "A golden retriever running through a park" is a scene description. "A low-angle tracking shot following a golden retriever sprinting through a sunlit park, camera moving at the dog's running speed, shallow depth of field" is a shot description, and it's the difference between a generic clip and one that looks intentional.
Section 02: Camera Language Actually Works
Sora responds well to real cinematography vocabulary: tracking shot, dolly zoom, aerial shot, handheld, static wide shot, close-up. These terms carry specific, learned visual meaning, and using them gets you much closer to a specific look than trying to describe camera movement in plain, non-technical language.
💡
Pro Tip: Specify camera movement AND speed. "Slow tracking shot" and "fast tracking shot" produce meaningfully different energy in the final clip, don't leave pacing to chance.
Section 03: Lighting and Time of Day Set the Mood
Lighting does more to establish mood than almost any other single word choice. "Golden hour," "overcast diffused light," "harsh midday sun," and "blue hour twilight" each tell a completely different visual story even with an identical subject and action. Pair your lighting description with the emotional tone you're going for, warm and nostalgic calls for golden hour, tense and moody calls for harsh shadows or low light.
Section 04: Describe Motion With Precision
Because this is video, not a still image, motion deserves its own careful description. Vague motion ("something is happening in the background") produces unpredictable, often distracting results. Specific motion ("steam rising slowly from a coffee cup in the foreground while people walk past blurred in the background") gives the model concrete physics to render.
⚠️
Watch out: Overloading a single prompt with too many simultaneous actions tends to produce a muddled result. If your shot needs three distinct things happening, consider whether it's really two or three separate shots edited together instead of one.
Section 05: Iterate on One Element at a Time
Just like image generation, changing multiple variables between attempts (camera angle, lighting, and subject action all at once) makes it impossible to tell what actually improved or hurt the result. Lock everything else and change one variable per generation: first nail the camera movement, then adjust lighting, then refine the action.
Key Takeaways
Structure prompts as shots: subject, action, camera, setting, in that order.
Use real cinematography terms, tracking, dolly, aerial, they carry specific visual meaning.
Lighting and time of day set the emotional register of the clip; choose them deliberately.
Describe motion with the same precision you'd use for a still subject.
Change one variable per iteration so you can tell what's actually working.