Free guide · no signup

The AI Video Prompting Guide

Every major image and video model wants to be prompted differently. This is what each one actually expects — side by side, with the failure each one produces and how to fix it.

Seedance takes audio as delimiters inside the prompt. Kling wants a master prompt plus per-shot blocks. MiniMax Hailuo takes literal camera commands in square brackets. Veo burns subtitles into the video unless you explicitly tell it not to. FLUX has no negative prompts at all.

Every model maker documents their own model and nothing else, because no vendor has a reason to write about a competitor's syntax. So this puts them in one place: ~15 image and video models, the structure each one wants, the audio and dialogue formats, the camera controls, and what actually transfers between them.

How to read the tags. A = officially documented by the platform or the model's maker. D = community-reported — useful, but unverified. Sources are listed at the end, and the honest gaps are marked as gaps rather than filled with guesses.

Everything here is dated 2026-08-10. These platforms ship changes weekly. When the app disagrees with this guide, the app is right.

Feeding this to an AI assistant? That's what it was built for. Give the whole page to Claude or ChatGPT and tell it to use this as its reference when helping you write prompts — it'll pick the right model and the right syntax instead of guessing.

How to use it: apply the doctrine below to structure your direction, pick the model with the model map, write the prompt using that model's sheet, and when it fails, use that model's failure-fix notes — changing one thing per retry.

Part 1 — Universal doctrine

These hold on every model I looked at. Learn these and most of the per-model detail becomes adjustment rather than relearning.

  1. One clip = one continuous shot. Never ask for cuts unless you're in an explicit multi-shot mode (Kling 6-shot, Veo timestamps, Wan 2.7 "Shot n", Seedance numbered shots). Asking for cuts inside a normal generation produces mush.
  2. Subject first, style last. Lead with who and what exists and who owns each object, then the action in time order, then camera, then environment, then style. Front-loaded words get more weight in nearly every model.
  3. The image-to-video golden rule: never re-describe the start image. The image already decides who is present, the framing and the light. Prompt only what moves, what causes it, the camera, the audio, and where it settles. This one changes more than any other.
  4. One primary camera move per clip. Stacked moves cause warping and drift. Name the move in film-school verbs — push-in, dolly, orbit, tilt — and state what the camera is NOT doing to lock it: static shot, no cuts, no zoom.
  5. Timed beats are the universal pacing control. 0–3s: [beat]. 3–6s: [beat]. Works on Seedance, Veo, Wan 2.7, Hailuo H3 and Cinema Studio, and it's badly underused.
  6. Positive phrasing beats negatives. Say "empty street," not "no cars." Where negatives are supported, list concrete artifacts ("frozen lips, warping fingers, extra limbs") — never vibes.
  7. Identity discipline. Write one authoritative character description — face, age, build, hair, wardrobe, distinguishing marks — and reuse it verbatim in every prompt. Never write "keep consistent"; enumerate the features. Anchor with a trained character ID where the platform has one, and tag the reference in every prompt.
  8. Every reference image gets a job. @image1 provides the character's face only. Do not use its background. Unassigned references leak unwanted content into the frame.
  9. Dialogue rules. Label speaker and tone, put action before the line, keep lines one-breath short (≈2 words per second of clip), and write them verbatim. Suppress unwanted audio explicitly — "No dialogue. No background music" — because silence left unstated usually gets scored.
  10. Keep talking-head clips under ~30s (gesture repetition sets in beyond that); under 20s hides eye-behaviour artifacts. A
  11. Iterate by targeted edit, not re-roll, wherever it's supported — Seedance region edit, Kling scene editing, GPT Image 2 change/preserve. One variable per retry; freeze what worked.
  12. Never rely on exact on-screen rendered text in video. Use subtitle syntax where it exists, or add the text in post.
  13. Repair method: observe before diagnosing. "Six fingers on the right hand" is an observation; "the model is bad at hands" is a story. Find the first failing moment — the ranges before and after it are often still usable trimmed.

Part 2 — Model map

Which model for which job.

JobFirst choiceAlternatives
Character design stills + consistencySoul 2.0 + Soul IDNano Banana Pro (14 refs), Seedream
Cinematic keyframes (photographic look)FLUX.2 Pro / Nano Banana ProSeedream 4.5/5.0, GPT Image 2
Storyboards / multi-frame sequencesPopcornNano Banana
Instruction-style image editsGPT Image 2Nano Banana
Text in images (titles, signs)GPT Image 2 / Nano Banana ProRecraft (typographic systems), Seedream 4.5
Hero video shots (long, audio, refs)Seedance 2.5 (30s, 50 refs, same-pass audio)Seedance 2.0
Cheap drafts / B-roll volumeKling 2.6 / MiniMax HailuoKling 3.0
Multi-shot sequences in one generationKling 3.0 (6 cuts)Veo timestamps, Wan 2.7
Start-frame → end-frame continuityKling 3.0 / WanHailuo 02
Atmosphere/realism + native soundscapeVeo 3.1Seedance 2.5
Dialogue / lipsyncKling 3.0 Omni (Voice Binding)Veo 3.1, Wan Speak, LipSync Studio
Transformations / elemental FX / animeHailuo 2.3
Camera-move spectacle on a stillDOP presets (65 moves)Kling
Final polishUpscale (Topaz; Iris for faces)

Part 3 — Image models

Soul 2.0 + Soul ID

Character-consistency backbone

  • Preset-driven, not prompt-driven A: pick one of 20+ aesthetic presets (Mystique City, Warm Ambient, Editorial Street Style, Nature Light, Theatrical Light…) — the preset carries lighting, camera and style. Your prompt supplies subject, action, setting, outfit. It understands fashion and editorial language natively.
  • Soul ID training A: upload 20–80 photos of the person — well lit, varied angles and expressions, at least one full-height shot, from the last 4–5 months. Avoid sunglasses, heavy shadows, crops and group shots. Quality beats quantity. Trains in about 3–5 minutes.
  • Cross-model bridge A: the trained character appears under Elements — type @ in the prompt field of Nano Banana, Seedream, Kling, Seedance or Cinema Studio and select the character. This is how one face survives a whole film.
  • Don't re-describe the face in the prompt when the character is selected — the ID handles it.
  • Failure modes: extreme style shifts cause small drift (expected); a weak training set gives a weak lock; one identity per Soul ID, so multi-character scenes need multiple Elements.
[Character selected] + [preset] + "wearing X, doing Y, in Z location"

Popcorn

Storyboards

  • Manual mode is the filmmaker mode A: per-frame prompts — set shot type (close-up, wide, over-the-shoulder), action and lighting per cell. Auto mode takes one world-prompt and returns a 4/6/8-frame narrative arc.
  • Up to 4 reference images, roles assigned by ordinal in prose: "The man from image one standing in the forest from image two, holding the object from image three."
  • Subject first — the first-named subject becomes the tracked subject. Concrete action verbs. Lighting and camera early.
  • Face, clothing and proportions are tracked across frames automatically; you can fix a single frame without breaking the rest.
  • Failure modes: low-res references degrade every frame; mismatched aspect ratios; abstract prompts.
Frame N: [shot size + angle] of [subject from image 1] [action] in [setting from image 2], [lighting], [mood]

Nano Banana / 2 / Pro

Google — reference control and text rendering

  • Narrative prose, not tag lists A: [Subject] + [Action] + [Location] + [Composition] + [Style]. Positive framing only. Edits start with a strong verb.
  • Up to 14 reference images — the largest reference budget anywhere here. Assign roles in prose ("the attached sketch as structure, the attached fabric as texture").
  • Responds strongly to photographic vocabulary A: camera hardware ("shot on Fujifilm"), lens and aperture ("low-angle, shallow depth of field f/1.8"), lighting design ("chiaroscuro," "golden-hour backlighting"), film stock ("1980s colour film, slightly grainy"). Name materials precisely — "navy blue tweed," not "nice fabric."
  • Text rendering is a headline strength (especially Pro): put literal words in quotes, name a font style per line, multilingual is fine.
  • Edits: conversational instructions plus an explicit "keep the rest identical." Iterate in follow-ups rather than one giant prompt.
  • 1K/2K/4K; ratios 21:9 → 9:16. Soul ID attaches via @Elements.

Seedream 4.5 / 5.0

ByteDance

  • 4.5: a five-part order — subject → style → composition → lighting → technical. Front-load the character description. Sweet spot is 30–100 words; under 15 goes generic, over 150 and the instructions start competing. D
  • 5.0: six elements (format, subject, composition, lighting, in-image text, style). Order-tolerant but completeness-rewarded; stay under ~600 words. Up to 10 references, named by content ("the tan leather-bound journal") rather than "image 3." Pin the invariants: "Keep the exact silhouette, the exact camera angle, the exact studio lighting from the reference."
  • Very responsive to lighting vocabulary (golden hour, low-key, dramatic side light) and camera terms ("85mm lens," "rule of thirds").
  • Text: 4.5 renders it well; 5.0 garbles long text — acknowledged — so keep 5.0 text short.
  • Failure fixes: inconsistent subject → lead with the subject and add specific traits; style miss → explicit "not cartoon-like" plus reinforcing synonyms. Watch for 5.0's over-smooth skin and action overshoot on dynamic poses.

FLUX.2 Pro

Black Forest Labs

  • Subject + Action + Style + Context, most important first. 30–80 words ideal. A
  • Two unique powers. Literal JSON prompting for multi-element precision — {scene, subjects:[{description, position, action}], style, color_palette:["#hex"], lighting, camera:{lens, angle, depth_of_field}} — and hex-colour binding: "The sofa is deep teal hex #1B6B6F" holds far better than naming a colour.
  • Multi-reference 8–10 images, roles by index ("clothing from image 1, style from image 2").
  • Loves hardware realism: "Shot on Hasselblad X2D, 80mm, f/2.8," and era emulation ("2000s digicam").
  • No negative prompts. It's guidance-distilled — write "sharp focus throughout," never "no blur." This is the single most common mistake people bring over from SD-era habits.

GPT Image 2

OpenAI

  • Order: background/scene → subject → key details → constraints. Labelled segments beat dense paragraphs. A
  • The core skill is change/preserve editing: Change: [exactly what]. Preserve: [face, identity, pose, lighting, framing, background]. Constraints: [no extra objects]. Restate the preserve list on every iteration — drift accumulates silently.
  • Best-in-class small text (quotes plus typography; spell hard words letter by letter).
  • Photorealism: say "photorealistic," add real-texture cues ("pores, fabric wear"), and prefer "35mm film photograph" over polish words.
  • Character template: "Same outfit, features and proportions; new scene and action. Do not redesign the character" plus an anchor image.

Recraft

Titles, posters, prop graphics

  • Structured order: concept → environment → framing/pose → attributes → spatial relations → lighting → camera → mood. A
  • Typographic systems are the specialty: quoted text plus format type, hierarchy, placement and colour blocking. For vector and logo work: "no gradients, no shadows," strict palette.
  • Avoid stacked evaluative adjectives — "stunning," "gorgeous" — they measurably degrade output. Style-reference copies layout, angle and composition strictly, which is what makes uniform title-card sets possible.

Draw-to-Video / Draw-to-Edit

Input mode, not a model

  • Sketch the motion directly on the start frame A: arrows for direction and speed (longer = stronger, curved = arcs), short text labels for actions ("man walks"), numbers for sequence order.
  • Draw-to-Edit takes up to 8 references — drop an object in and size it to match perspective.
  • One redraw beats five prompt rewrites.

Upscale

Finishing

  • Image presets: Standard, High Fidelity v2, Text Refine (signs and titles), Art & CG, Low Res. Video: Proteus (standard), Iris (faces and portraits — use it for character shots), Gaia, Theia. A
  • Settings: Face Enhancement 70–85%; denoise ≤50; sharpness above 70 causes artifacts. AD
  • Upscaling will not fix composition or heavy motion blur. Fix those upstream.

Part 4 — Video models

Seedance 2.5 / 2.0

ByteDance — the workhorse

  • Official formula A: Subject + Action + Setting + Visual style + Camera work + Sound — only subject and action are mandatory. For multi-shot, declare shot count, total duration and aspect ratio upfront, then baseline (film stock, aesthetic) → environment and character → numbered shots.
  • 2.5 A: up to 30s per generation. The UI selectors (era, genre, lighting, physics, lens, tone, pacing) do heavy lifting, so the prompt can stay short — around 350 words for a 30s clip is the official sample scale. Up to 50 references (practically ~8–12 image subjects). Audio is generated in the same pass. Region edit fixes one object, face or background without re-rolling the clip — use it instead of retrying.
  • References: "@image1 provides the character's face. Do not use its background." For multi-angle references of one object: "all three images define one single [object]."
  • Camera: name the real move, and state exclusions to lock POV — "No cuts, no zoom, natural head movement." Handheld: "unstabilized, micro-jitters, abrupt jerks."
  • Audio AD: specify a sound list or "NO MUSIC / NO SFX" — silence left unstated gets a score. 2.5 delimiters D: music ( ), SFX < >, dialogue { }, subtitles 【 】. Declare language and accent before the lines.
  • Failure fixes: drift on long clips → timed beats plus per-segment camera; feature dilution → fewer references; background leakage → explicit exclusions; random cuts → "no cuts."
@image is the first frame. 0–3s: [beat]. 3–8s: [beat]. Camera: slow push-in only. Preserve natural skin from reference. Audio: [ambience], (soft score). No cuts.

Kling 3.0 / 2.6

Kuaishou — multi-shot, cheap drafts, lipsync

  • Write directions to a scene, not a list of objects. D The pattern is a master prompt (style anchor, 2–3 traits per character, environment, palette) plus per-shot prompts (framing, camera, beats, dialogue, duration).
  • Multi-shot AD: up to 6 shots / 15s in one generation. Label each shot with framing and duration; characters and grade carry across automatically.
  • Start and end frame supported (or end-frame only) — the continuity power tool between scenes.
  • Dialogue D: [Name: role, tone]: "Line." Action before the line; unique names, never pronouns; "Immediately" between speakers kills dead air. Line budget is roughly 8–12 words per 5s before lip sync falls apart. Voice Binding (3.0 Omni, official) binds a voice from 5–30s of clean audio or a 3–8s single-speaker video.
  • Negatives: list concrete artifacts — "frozen lips, warping fingers, jittery eyes, character drift."
  • Kling 2.6 is the budget single-shot B-roll and cheapest lipsync option; use the simple subject + action + environment + camera formula there.
MASTER: [style], [Char A: desc], [environment], [palette]. SHOT 1 (4s): medium two-shot, slow push-in; A pours coffee; [A: warm]: "line." SHOT 2 (3s): reverse close-up; Immediately [B: quiet]: "line." NEGATIVE: frozen lips, character drift.

Veo 3 / 3.1

Google — atmosphere and native audio

  • Formula A: [Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance], written in prose. Clips 4/6/8s, 16:9 or 9:16.
  • Image-to-video: the image carries subject and scene; the prompt carries motion and audio. Ingredients-to-Video supplies references by role ("Using the provided images for the detective and the office…"). First+last frame: describe the camera path bridging the two stills.
  • Audio syntax A: dialogue in quotes with delivery in prose (She says, "We have to leave now," in a weary voice); SFX: thunder cracks; Ambient noise: quiet hum of the bridge; music as a described line ("a swelling orchestral score begins").
  • Timestamp prompting gives you mini-edits in one generation: [00:00-00:02] Medium shot… [00:02-00:04] Reverse shot…
  • The gotcha: quoted dialogue tends to burn in captions. Append "no subtitles, no captions, no on-screen text." D Describe exclusions positively where you can — "desolate landscape with no buildings" style.

Wan 2.5 / 2.6 / 2.7

Alibaba — frame control and Speak

  • Formula A: Entity + Scene + Motion (plus aesthetic control, stylization and sound).
  • Image-to-video: the image owns entity, scene and style; the prompt is motion and camera only, with tempo words ("slowly," "quickly"). Re-describing the image makes it fight you.
  • Camera: verb phrases (push-in, pull-out, tracking, tilt up, fixed camera). Keep orbit arcs under ~45°. "Generate single shot" prevents unwanted cuts.
  • Audio A: voice = line + emotion + tone + speed + timbre + accent; speaker-labelled lines (Explorer: "…") with stage directions ("trembling voice, calling out over wind"). Suppress with "No dialogue" / "No background music." Wan Speak (LipSync Studio) maps speech onto footage while preserving identity — the fix when word-exact lipsync matters.
  • 2.7 multi-shot: Shot 1 [0–3 s]: …, references cited as "Image 1" / "Video 1". 2.6 takes up to 3 reference characters ("character1…").
  • Known limits A: real people's names are rejected; one clip is one continuous shot outside explicit multi-shot; exact on-screen text won't render; long choreography breaks, so keep actions simple.

MiniMax Hailuo 02 / 2.3 / H3

Bracket director syntax and physics

  • Hailuo 02 A: the unique bracket director syntax[Push in], [Truck left], [Pedestal up], [Static shot], [Tracking shot]… up to three combined: [Truck left, Pan right, Zoom in]. More than three degrades. First and last frame supported.
  • Hailuo 2.3 A: short, physical prompts. Verbs that translate to physics — "dissolve," "ignite," "shift into light." Strengths are transformations, elemental FX and anime line stability. Camera comes from preset categories in the UI, not inline. Image-to-video rule: describe only what changes, and never ask for a wide shot from a close-up input. Fixes: morphing → one action only; static output → add motion verbs; "deep-fried" look → delete "8k, masterpiece" quality tags. No native audio documented.
  • H3 Dno official prompt documentation exists yet. Community formula: reference description + core idea (1–2 sentences) + a time-stamped visual process; 50–120 words typical; camera in natural language tied to time windows ("0–3s: camera slowly pushes in"). Write dialogue verbatim — paraphrase leaves no lipsync reference. Specify the ending state or the last frame is unusable. For identity, enumerate every feature; never "keep consistent."

The 65 camera presets (DOP)

Named moves as selectable controls

Zooms: Zoom / Rapid Zoom / Crash Zoom In & Out, YoYo Zoom, Eating Zoom, Eyes In, Mouth In · Dolly: Dolly In/Out/Left/Right, Super Dolly, Double Dolly, Dolly Zoom (the Vertigo warp) · Orbit: 360 Orbit, Arc L/R, Lazy Susan, 3D Rotation · Crane & aerial: Crane Up/Down/Over-The-Head, Jib, Aerial Pullback, Overhead, FPV Drone, Flying Cam · Pan & tilt: Pan L/R, Whip Pan, Tilt, Dutch Angle, Incline · Handheld & rig: Handheld, Static, Wiggle, Snorricam (body-locked), Robo Arm, BTS, Hero Cam, Glam, Head Tracking, Object POV · Vehicle: Car Chasing, Car Grip, Buckle Up, Road Rush · FX: Through Object In/Out, Bullet Time, Focus Change, Fisheye, Low Shutter · Time: Hyperlapse, Timelapse variants.

The technique that actually works D: the preset sets the motion engine, then repeat the move by name in the prompt text to reinforce it. Dual-naming is the most repeatable approach — better than either alone. One primary move per ~5s clip. Order: camera move → subject (concrete) → action → setting → lighting → lens/style → mood. When animating a still, add "preserve the original face, lighting, and geometry."

Cinema Studio

The assembly environment

  • Layered prompting A: (1) scene setup — location, time, weather; (2) numbered shot breakdown with camera type per shot ("Shot 1: telephoto… Shot 2: handheld tracking…"); (3) character action, blocking and dialogue; (4) technical specs — lens, movement, lighting quality.
  • Hardware selectors: 5 camera profiles (ARRI/RED/IMAX), 11 lenses from 8–75mm. Focal-length choice, not adjectives, is what makes it cinematic — telephoto compresses, handheld shakes, wide gives context. D
  • Lighting written contextually ("dusk sky," "warm tungsten," "silhouetted by sunset"), not as parameters. Pacing inline ("0–4 seconds: …").
  • References via @image / @video tags in every prompt. Character drift → character sheets plus tagged references every time. One camera move per shot.

Part 5 — The consistent-character pipeline

Putting all of it together, in the order that costs least.

  1. Train a character ID for each main character (20–80 recent varied photos, at least one full-body) — or design the character in stills first, then train from the approved set.
  2. Design keyframes as stills (FLUX.2, Nano Banana, Seedream, character via @Elements). Iterate identity, composition and light at image prices before paying video prices. This is the step most people skip and most regret skipping.
  3. Animate image-to-video — Seedance 2.5 for hero shots, Kling or Hailuo for drafts. Prompt only motion, camera and audio. Use start/end frames (Kling, Wan) where shots must connect.
  4. Dialogue through Kling Omni Voice Binding, Wan Speak or LipSync Studio. Lines short, speaker-labelled, clips under 30s.
  5. Fix, don't re-roll — Seedance region edit, Kling scene edit, one variable per retry.
  6. Finish: upscale (Iris for faces), then assemble.

Part 6 — Sources

Where to go deeper, and what I could not verify.

Official platform documentation

Model chooser, Seedance 2.0 and 2.5 prompting guides, the Kling 3.0 user guide, Wan 2.5 text-to-audio, the Hailuo 2.3 creative guide, the lipsync guide, Cinema Studio 3.0, the camera-controls page, Soul ID creation, Popcorn storyboards, Draw-to-Edit, and the Upscale playbook — all published on higgsfield.ai under /blog, /creator-hub/help-center and /camera-controls.

Official model-maker documentation

Veo 3.1 and Nano Banana — the ultimate prompting guides on cloud.google.com/blog/products/ai-machine-learning · Wanalibabacloud.com/help/en/model-studio/text-to-video-prompt · FLUX.2docs.bfl.ml/guides/prompting_guide_flux2 · GPT Image — the image-gen prompting guide in the OpenAI cookbook · Recraftrecraft.mintlify.app/prompt-engineering-guide · Kling Omnikling.ai/quickstart

Study material with real prompts

Higgsfield open-sourced the full prompts for their films Hell Grind, Zephyr and Mork, with a 19-minute breakdown video — the single best study resource available, because it's real production prompts rather than tutorial examples. Contest-winner workflow breakdowns are published at contest-workflows.higgsfield.app.

Community references

Kling 3.0 guides on blog.fal.ai and videoai.me · Seedance formula on tryonr.com · Seedream on fal.ai/learn and runware.ai/docs · Hailuo on akool.com · H3 on minimax3.com · camera cheat sheet on memons.ai.

Known gaps — stated, not filled with guesses

No official MiniMax H3 prompt documentation exists yet; all H3 guidance here is community-sourced.

ByteDance publishes no first-party English Seedream guide — host documentation is the best available.

Seedance 2.5's bracket delimiter audio syntax is community-reported. Verify it in the app.

No official doctrine exists for combining camera presets with prompt text — dual-naming is community practice that works, not documented behaviour.

About this guide

Published free by ZEROFAYYZ. Written by a working creator, not a vendor. Nothing here is sponsored by or affiliated with any platform or model maker, and all product names and trademarks belong to their respective owners.

Share it. Copy it, quote it, repost it, feed it to your AI. Attribution appreciated, not required.

Platform facts go stale fast. If you find something that's changed, say so — corrections make the next version better.