How to prompt Kling 3.0: camera control that works

By the Infer teamUpdated

Kling 3.0 Pro's product page claims frame-by-frame camera direction and ranks it "Top 1" for camera control among the models Infer tracks (Kling 3.0 Pro on Infer). The one thing that matters most when you're writing prompts for it: treat the camera move as its own clause, named and directional, the same way you'd separate the subject from the setting. Everything below is a set of templates built from that documented control surface, adapted to your own subject rather than copied verbatim: we haven't rendered these, and we're not going to pretend we have.

Kuaishou's own spec sheet also lists native 4K capture (the model ships output at 1080p) and "best in class" stylized motion, which is why the templates below lean toward choreography, anime, and camera-driven brand work, the use cases the documentation itself points at (Kling 3.0 Pro on Infer).

The prompt anatomy that works

Build every Kling 3.0 prompt from four clauses, in this order:

  1. Subject + action: who or what, doing what, in plain language.
  2. Camera move: one named verb (orbit, dolly, crane, tracking, static) with a direction and a speed: this is the clause the documentation says the model is built around.
  3. Style: the register the shot lives in (photoreal, anime, stylized/graded, music-video), since Kling's own use-case list spans "stylized music videos" through "anime/cinematic shorts."
  4. Aspect ratio: 16:9, 9:16, or 1:1, decided before you write the rest, since a 9:16 dance clip and a 16:9 hero shot want different framing language in clause one.

A labeled example:

Subject/action: a violinist performs alone on an empty stage, bow arm sweeping through a long note
Camera move: slow orbit, clockwise, quarter turn, medium shot
Style: photoreal, warm stage lighting
Aspect ratio: 16:9

Drop the camera clause and you're asking the model to fill in a move on its own, when the documented control exists specifically so you don't have to. The four-clause structure is the whole trick; the templates below apply it at increasing difficulty.

Nine templates, easy to hard

Each is a starting point to adapt, not a script with a guaranteed result.

1. Static establishing shot

Subject/action: a lighthouse stands on a rocky point at dusk, waves breaking below
Camera move: static, locked-off, wide
Style: photoreal, cool blue hour light
Aspect ratio: 16:9

Why start here: no camera move to get wrong. It isolates style and lighting phrasing as the only variables before you add motion. Try this template on Infer →

2. Tracking shot

Subject/action: a cyclist rides down a tree-lined street, morning light flickering through leaves
Camera move: tracking, lateral, matched to the cyclist's speed, medium shot
Style: photoreal, dappled light
Aspect ratio: 16:9

Kling's documented camera vocabulary treats tracking as a named, distinct move from a dolly or a pan, so say the one you mean rather than a generic "follow the subject."

3. Choreography and dance

Subject/action: a dancer executes a full turn and drops into a low extension, one continuous phrase
Camera move: crane down, following the drop, medium to wide
Style: stylized, high contrast stage lighting
Aspect ratio: 9:16

Dance and choreographed motion are named use cases on Kling 3.0's own page, so match the crane's speed to the turn in your written description, since the camera clause and the subject's motion are the two things the model has to reconcile in the same shot.

4. Anime-style character beat

Subject/action: an anime-style swordsman sheathes a blade after a duel, wind moving through a cloak
Camera move: slow push in, low angle, medium close-up
Style: anime, cel-shaded, high contrast rim light
Aspect ratio: 16:9

"Anime/cinematic shorts" is on Kling 3.0's documented use-case list alongside stylized motion: the style clause is doing real work here, separate from the camera clause, so keep them as two distinct lines rather than folding "anime camera" into one phrase.

5. Multi-shot brand spot

Subject/action: a watch face catches light as a hand adjusts the strap
Camera move: 0-4s static macro close-up → 4-8s slow orbit, quarter turn
Style: photoreal, studio lighting, deep shadow
Aspect ratio: 1:1

Splitting the camera clause into timestamped segments is the documented way to ask a single generation for more than one move: one continuous take, two named camera beats in sequence.

6. Image-to-video from a product still

Reference image: [product hero shot, watch on a dark surface, provided]
Subject/action: the same watch, seconds hand ticking, background stays still
Camera move: static, locked-off, close-up
Style: match the reference image's lighting
Aspect ratio: 1:1

Kling 3.0's catalog listing is text/image-to-video, so this reference-anchored mode is a documented path, not a workaround: the still frame sets the opening composition, and the rest of the prompt governs what happens after it.

7. Multi-reference scene (5 images)

References: [1: character face, 2: costume, 3: location, 4: prop, 5: lighting reference]
Subject/action: the character from reference 2 walks into the location from reference 3, carrying the prop from reference 4
Camera move: dolly follow, medium, eye level
Style: match reference 5
Aspect ratio: 16:9

Infer's changelog dated July 20, 2026 expanded reference-image support to up to five images on Kling 3.0 (Infer changelog). Give each reference exactly one job (face, costume, location, prop, lighting) rather than letting two images both define the subject's appearance.

8. Vertical social cut

Subject/action: a street food vendor flips a skewer over open flame
Camera move: handheld-style tracking, close, fast
Style: photoreal, high contrast, warm practical light
Aspect ratio: 9:16

9:16 is a documented aspect ratio on Kling 3.0, same as the 16:9 and 1:1 options, so write it into the anatomy explicitly rather than assuming a square or landscape prompt will crop cleanly to vertical.

9. Extended camera sequence

Subject/action: a rock climber ascends a sheer face at sunrise
Camera move: 0-5s static wide → 5-10s crane up, following the ascent, tightening to medium
Style: photoreal, backlit, long shadows
Aspect ratio: 16:9

The hardest template on this list: two camera moves, a full 10-second duration, and a subject whose motion has to match the crane's timing throughout. This is where the frame-by-frame camera claim is doing the most documented work, and where a first pass is least likely to land exactly as written.

Try Kling 3.0 Pro in the Infer playground →

What the docs don't cover

Two areas Kling 3.0's own documentation is silent on, worth planning around rather than guessing at.

Audio isn't part of this version. Infer's Kling 3.0 Pro page lists no native audio, unlike Seedance 2.0 Pro's joint audio-video generation or Veo 3.1 Fast's native sound (Kling 3.0 Pro on Infer). There's no documented in-model audio clause to write into a prompt here, so the working pattern is a silent Kling render plus a separate pass through a dedicated audio model: ElevenLabs Eleven v3 is hosted on Infer for voiceover, dialogue, or ambient tracks, including 30-second voice cloning and inline emotional tags.

Text-in-video isn't a claimed capability. Neither Kling 3.0's own product page nor Infer's catalog entry makes a legibility claim for on-screen text, signage, or captions the way Ideogram 3.0 or Nano Banana 2 do for still images. Treat that as unclaimed territory rather than an assumed strength: if a shot depends on readable text, that's a documented gap to design around (keep lettering out of frame, or composite it in post), not a prompt problem to try to solve.

We haven't run generations against either of these to confirm a specific failure shape, so we're not going to describe one. If you've hit a documented, evidence-backed community report on either front, it belongs here; this section gets updated as one surfaces.

Settings that matter

Duration and camera complexity are the two levers that change what a test costs. Kling 3.0 Pro clips run 5-10 seconds typically, at $0.10 per second flat on Infer: a 5-second iteration is $0.50, a 10-second one is $1.00. Requests are async and queued, with a default rate limit of 60 requests per minute, so batch your template variations rather than firing them one at a time and waiting.

Worth knowing before you commit to a template: Infer's own leaderboard page cites Kling 3.0 1080p Pro at #3 overall with a 1248 Elo score, from an Artificial Analysis snapshot dated 2026-04-30 (Infer leaderboards). Artificial Analysis's live arena page, pulled the same day this guide was compiled, ranks the same model #6 at 1111 Elo (Artificial Analysis video leaderboard). Both numbers are real, they're just from different dates on a leaderboard that reshuffles fast, so cite whichever snapshot you're quoting, not both blended together.

Reference images are the other setting to plan test passes around: up to five as of the July 20, 2026 changelog update. Full worked pricing, including a same-clip comparison against running Kling through a reseller, is in the Kling 3.0 Pro pricing guide.

Infer hosts every model referenced in this guide, Kling 3.0 Pro included, so nothing here is written to make the camera-control claim look better than the documented gaps around audio and text warrant.

Test your next prompt with Kling 3.0 Pro on Infer →

Related reading: the guides hub, how to prompt Seedance 2.0, Kling 3.0 Pro vs Veo 3.1 Fast, Kling 3.0 Pro vs Seedance 2.0 Pro, Wan vs Kling vs Hailuo, and image-to-video models if the reference-image workflow above is the part you came for.

Frequently asked questions

What camera-move keywords does Kling 3.0 recognize?

Kling's own product page claims frame-by-frame camera direction and a 'Top 1' ranking for camera control, so name the move the way a shot list would: orbit, dolly, crane, tracking, static, plus a direction and a speed (Kling 3.0 Pro on Infer). A prompting guide from fal.ai, which also hosts the model, lists profile shots, macro close-ups, tracking shots, POV, and shot-reverse-shot dialogue as camera language Kling responds to . Vague phrasing like 'nice camera work' isn't documented as producing anything specific.

Which aspect ratios does Kling 3.0 support?

16:9, 9:16, and 1:1, per Kling 3.0 Pro's documented specs on Infer (Kling 3.0 Pro on Infer). Pick the ratio before you write the shot: a 9:16 dance clip and a 16:9 establishing shot call for different framing language in the subject/action clause, not just a different export size.

Does Kling 3.0 generate audio, and what's the workaround?

Not on the Pro tier hosted on Infer — the model page lists no native audio for this version (Kling 3.0 Pro on Infer). The documented workaround is to generate the clip silent and add dialogue, ambient sound, or voiceover afterward with a dedicated audio model; ElevenLabs Eleven v3 is hosted on Infer for that pass, including voice cloning and inline emotional tags.

How many reference images can Kling 3.0 use?

Up to five, as of Infer's changelog entry from July 20, 2026, which expanded reference-image support to Kling 3.0 and Seedance 2.0 together (Infer changelog). Give each reference one job — face, wardrobe, location, prop, lighting — rather than stacking several images that all try to define the same thing.

What does Kling 3.0 cost per prompt iteration?

$0.10 per second on Infer, so a 5-second test costs $0.50 and a 10-second one costs $1.00 (Kling 3.0 Pro on Infer). Full worked examples, including how that compares to running Kling through a reseller, are in the Kling 3.0 Pro pricing guide.

Sources

Related reading