How to prompt Seedance 2.0: camera, lighting, audio

By the Infer teamUpdated

Seedance 2.0 Pro is built for director-style prompts, not scene descriptions. The single thing that matters most: treat camera move, lighting direction, and audio intent as three separate clauses in every prompt, the same way a shot list separates them. That's because the model generates video, camera motion, and audio in one joint pass rather than compositing them afterward (Seedance 2.0 Pro on Infer).

That joint generation is also why sloppy prompts risk failing in a specific way here. Since Seedance's documented architecture generates camera, lighting, and audio in a single pass, a vague camera clause risks dragging the lighting and audio pass down with it rather than staying an isolated weak spot. Get specific early and the rest of the prompt has a better documented basis to hold together.

The prompt anatomy that works

Seedance 2.0 Pro's own product page documents "director-level control": named camera moves, lighting direction, and shadow density as first-class prompt inputs, not afterthoughts layered on top of a generic description (Seedance 2.0 Pro on Infer). Build every prompt from five clauses, in this order:

  1. Subject + action: who or what, doing what, in plain language.
  2. Camera move: one named verb (orbit, dolly, crane, static) with a direction and a speed.
  3. Lighting direction: where the key light sits relative to the subject (backlit, side-lit, top-down).
  4. Shadow density: how hard or soft the shadows fall (deep contrast vs. soft fill).
  5. Audio intent: what's audible and how it's mixed (dialogue, ambient, music, or silence).

A labeled example:

Subject/action: a barista pulls an espresso shot, steam rising past their face
Camera move: slow dolly in, eye level, medium to close-up
Lighting: backlit through a window, warm morning key
Shadow density: soft, low contrast
Audio: espresso machine hiss, cup clink, no dialogue

Drop any one clause and Seedance 2.0's documented behavior is to fall back to an undirected default for it, rather than the intentional camera, lighting, or audio choice you'd otherwise control. The five-clause structure is the whole trick; everything below is that structure applied at increasing difficulty.

Ten worked examples, easy to hard

Each example builds on the one before it, as a template you adapt rather than a script to copy verbatim.

1. Static product shot

Subject/action: a matte black wireless earbuds case sits closed on a marble surface, not moving
Camera move: static, locked-off, eye level, close-up
Lighting: soft top-down key, single source
Shadow density: soft, minimal
Audio: silence

Why this template works: no camera move to get wrong, no dialogue to sync. It isolates lighting and shadow phrasing as the only variables in play, which is why it's the structure to start from before adding motion. Try this template on Infer →

2. Dolly reveal

Subject/action: the same case opens, earbuds lit from inside the case
Camera move: slow dolly in, eye level, close-up to extreme close-up
Lighting: backlit from the case interior, cool key
Shadow density: deep contrast against the marble
Audio: soft mechanical click on open, then silence

Variation: swap "slow dolly in" for "orbit clockwise, quarter turn" to see the same product from a different angle. Seedance's documented camera vocabulary treats orbit and dolly as distinct named moves, not interchangeable synonyms, so naming the one you mean is the point of this template.

3. Choreographed scene with Foley

Subject/action: two dancers cross paths in a rehearsal studio, one spins under the other's raised arm
Camera move: crane up, following the spin, wide to medium
Lighting: side-lit, hard key from stage-left
Shadow density: deep, long shadows across the floor
Audio: footsteps on wood floor, fabric rustle, no music

This is the first template where camera and subject motion have to agree in timing. Match the crane's speed to the spin: a crane move that outruns the spin risks reading as the camera chasing the dancers rather than revealing them.

4. Talking character (lip-sync)

Subject/action: a woman in a kitchen explains a recipe step to camera, mid-sentence pause before "simmer"
Camera move: static, eye level, medium close-up
Lighting: soft window light, front-left
Shadow density: soft
Audio: dialogue: "...and then you let it simmer," warm, conversational pace, English

Seedance 2.0's documented lip-sync capability is built for a dialogue clause this explicit: full sentence, named pause point, named language (Infer's leaderboard notes "native audio + phoneme-level lip-sync" for this model). Because the audio and video come out of the same generation pass, a competing music or ambient clause in the same prompt risks interfering with that sync, so the template keeps dialogue-only prompts free of it until the read is locked.

5. Image-to-video via init_image

init_image: [product hero shot, earbuds case on marble, provided]
Subject/action: steam rises from an unseen espresso cup out of frame, camera holds on the case
Camera move: static, eye level, close-up
Lighting: match the source image's key
Shadow density: match source
Audio: ambient cafe murmur, low

init_image anchors the opening frame; everything else in the prompt still applies to what happens after frame one. Skip the "match the source image's key" lighting clause and you risk the color temperature drifting partway through the clip, since only the opening frame is anchored to the reference.

6. Multi-reference (5 images)

References: [1: character face, 2: wardrobe (red jacket), 3: street location, 4: prop (motorcycle), 5: time-of-day lighting reference]
Subject/action: the character in reference 2 walks toward the motorcycle in reference 4, in the location from reference 3
Camera move: dolly follow, medium, low angle
Lighting: match reference 5
Shadow density: hard, late-afternoon length
Audio: footsteps on pavement, distant traffic

Infer expanded reference-image support to up to five images on Seedance 2.0 as of the July 20, 2026 changelog entry. Give each reference exactly one job in the prompt (face, wardrobe, location, prop, lighting) rather than letting two references both imply the subject's appearance. Overlapping references risk a face that drifts halfway through the clip, since there's no documented way to tell the model which reference should win.

7. Ambient nature scene, no dialogue

Subject/action: a heron stands motionless in shallow water, then strikes at a fish
Camera move: static, then a fast push in timed to the strike
Lighting: overcast, flat, no hard key
Shadow density: minimal
Audio: water lapping, wind in reeds, splash timed to the strike

The strike and the push-in need to land in the same beat. Naming the timing relationship explicitly ("timed to the strike") rather than listing the two events separately is the structural way to ask for that sync.

8. Music-video cut with lighting changes

Subject/action: a singer performs a chorus, spins once on the beat
Camera move: orbit counter-clockwise, quarter turn, medium
Lighting: cuts from cool blue key to warm red key on the spin
Shadow density: deep contrast throughout
Audio: music: mid-tempo pop chorus, no dialogue

Tying a lighting change to an action beat ("on the spin") is how you specify a scripted cue instead of leaving the timing to chance. It's the same timing-language structure as example 7, applied to light instead of camera.

9. Timeline-notation sequence

Subject/action: a runner crests a hill at sunset
Camera move: 0-4s wide establishing shot, static → 4-8s slow push-in to medium → 8-12s orbit around subject
Lighting: warm backlit throughout, intensifying through the sequence
Shadow density: long shadows, deepening toward 12s
Audio: wind, footfalls, breathing, building

Splitting the camera clause into timestamped segments lets one generation carry three distinct moves in sequence rather than one move for the whole clip. That's the move once a single prompt needs to do more narrative work than one shot can hold.

10. Extreme close-up, hands doing precise work

Subject/action: a jeweler sets a small stone into a ring mount with tweezers
Camera move: static, extreme close-up, locked
Lighting: hard top-down spotlight
Shadow density: deep, tight
Audio: metal-on-metal tick, breathing, no dialogue

This is deliberately the hardest example on the list. Fine hand-and-finger work in extreme close-up is the category where independent testing has found the most breakage , so budget for a retry or two, and consider a medium shot as the fallback if the close-up keeps drifting.

Try Seedance 2.0 Pro in the Infer playground →

What doesn't work

No prompt trick fixes these. They're limits of the current model, not of your phrasing.

On-screen text still comes out garbled. Independent testing consistently reports that legible text (signage, screens, t-shirt logos, labels) renders as illegible shapes rather than real characters . The fix isn't a better prompt. Keep lettering out of frame, or add it in post rather than asking the model to render it.

Extreme close-ups on hands drift. As the template in example 10 is built to stress-test, fine finger work — gripping tools, precise gestures — is where independent testing reports extra or merged fingers most often . Wide and medium shots of hands are reported as reliable; the failure is specific to extreme close-up framing.

Vague camera language gets ignored, not interpreted. "Cinematic shot" and "dramatic angle" don't map to a specific move, so the model falls back to a static default . Name the verb: dolly, orbit, crane, push, pull, static.

Stacking multiple camera moves into one shot muddies the result. One move per shot is the working rule; asking for a dolly and an orbit in the same continuous take is a listed common mistake . Use timeline notation (example 9) instead, where each segment gets its own single move.

Settings that matter

Resolution and duration are the two levers with real cost consequences, and neither is free to get wrong. 1080p output rolled out on Seedance 2.0 Pro on July 13, 2026, alongside Veo 3.1 Fast, HappyHorse 1.1, and Wan 2.2 Flash, per Infer's changelog; before that, 720p was the ceiling. Clips run 5-10 seconds typically, up to 8-second clips at the outer end. At Infer's flat $0.13 per second, a 5-second experiment costs $0.65 and a 10-second one costs $1.30. There's no separate audio surcharge, since audio is generated in the same pass rather than billed as an add-on. Full pricing detail, including how Seedance 2.0 Pro compares to running the same model through fal.ai, is in the Seedance 2.0 Pro pricing guide.

Reference images are the other setting worth planning around: up to five as of the July 20, 2026 changelog update. Budget your test passes accordingly. If a multi-reference prompt isn't landing, dropping to three tightly-scoped references is a cheaper first fix to try than adding a sixth clause of description.

Infer hosts every model on this leaderboard, Seedance 2.0 Pro included, so nothing here is written to make the model look better than its failure modes warrant. The garbled-text and hand-drift limits above are as real as the director controls.

Test your next prompt with Seedance 2.0 Pro on Infer →

Related reading: the guides hub, Kling 3.0 Pro vs Seedance 2.0 Pro, Seedance 2.0 Pro vs Veo 3.1 Fast, Seedance 2.0 Pro vs Sora 2, the cheapest AI video generation APIs, and Sora 2 alternatives if you're migrating a pipeline that used Sora 2's joint audio-video approach.

Frequently asked questions

How do you prompt audio into a Seedance 2.0 clip?

State the audio intent as its own clause, the same way you'd write a camera move: name the sound source and its timing, e.g. "dialogue: warm, low register, mid-sentence pause before the last word" or "ambient: distant traffic, rain on glass, no music." Seedance 2.0 generates audio and video in one pass rather than stitching a separate track on afterward, so vague audio phrasing ("nice sound") gets ignored the same way vague camera phrasing does.

What camera-move keywords does Seedance 2.0 actually recognize?

Named film-language verbs work best: dolly in/out, crane up/down, orbit (with a direction, clockwise or counter-clockwise), push in, pull out, and static/locked. Pair each with a named speed (slow, fast) and a framing (wide, medium, close-up), one move per shot. Generic phrasing like "cinematic shot" or "dramatic angle" is treated as filler and produces a default composition.

How many reference images can Seedance 2.0 use?

Up to five, since Infer's changelog entry of July 20, 2026 expanded reference-image support to Seedance 2.0 and Kling 3.0 (previously fewer were supported). Multi-reference prompts work best when each image does one clear job: one for the subject's face, one for wardrobe, one for the environment, rather than five overlapping shots of the same thing.

What does the init_image parameter do, and does it cost extra?

init_image turns a text-to-video prompt into image-to-video: you supply a starting frame and Seedance 2.0 animates from it instead of generating the opening frame from scratch. It doesn't carry a separate line-item cost on Infer; you're still billed at the standard $0.13/second rate, only for generations that complete. Full worked pricing is in the Seedance 2.0 Pro pricing guide.

Sources

Related reading