Veo 3.1 API pricing, with worked examples
By the Infer teamUpdated
Veo 3.1 Fast costs $0.09-$0.10 per second of output on Infer, audio included . An 8-second clip, the model's native generation length, runs $0.72-$0.80. Google's own Gemini API charges $0.12/second at 1080p for the same Fast tier , and fal.ai charges $0.15/second once audio is turned on . Every number below is per second of generated video, not per API call.
What an 8-second clip actually costs
One native-length Veo 3.1 Fast clip, 1080p, with audio, priced across the three ways to run it:
| Infer | $0.09-$0.10/sec, audio included | $0.72-$0.80 |
| Google Gemini API (official, Fast tier, 1080p) | $0.12/sec, audio included by default | $0.96 |
| fal.ai (720p/1080p, with audio) | $0.15/sec | $1.20 |
Run Veo 3.1 Fast on Infer →. The same clip that costs $0.96 through Google's own API runs $0.72-$0.80 here, and there's no separate audio meter to budget around.
Chain four of those 8-second clips into a 32-second spot and the gap compounds: $2.88-$3.20 on Infer versus $3.84 direct through Google, versus $4.80 on fal.ai. Push the chain to Infer's documented ceiling of roughly 148 seconds and the same math puts Infer at $13.32-$14.80, Google's direct API at $17.76, and fal.ai at $22.20 , all still 1080p, all still with audio.
A vertical 9:16 cut of that same 8-second clip costs the same $0.72-$0.80 on Infer; the model doesn't meter by aspect ratio, though Infer's documentation notes vertical and square renders take about 10-15% longer to generate than 16:9 . That's a wait-time cost, not a billing one.
Price table across providers
| Infer | $0.09-$0.10/sec | Included, no surcharge | 1080p |
| Google Gemini API (official, Fast tier) | $0.10/sec (720p), $0.12/sec (1080p), $0.30/sec (4K) | Included by default at every tier | 4K |
| fal.ai | $0.10/sec (no audio) or $0.15/sec (with audio) at 720p/1080p; $0.30/sec (no audio) or $0.35/sec (with audio) at 4K | $0.05/sec surcharge for audio | 4K |
Google's own standard (non-Fast) Veo 3.1 tier runs $0.40/sec at 720p/1080p and $0.60/sec at 4K , roughly 3-4x the Fast tier this page covers. If you've seen a figure of "~$0.15/second" for Veo 3.1 Fast direct from Google, that's an aggregator estimate that doesn't match Google's own pricing page: the real rate splits by resolution at $0.10 (720p) and $0.12 (1080p), only reaching $0.30/sec at 4K. Use the resolution-specific numbers above, not the flat estimate.
One structural difference worth flagging: fal.ai is the only one of the three that meters audio separately. Infer's rate card lists no audio surcharge, and Google's own pricing page states a single combined rate described as the "video with audio price (default)" rather than splitting silent and audio columns. Don't assume fal's $0.05/sec audio premium applies elsewhere: the pages themselves say otherwise.
What affects the price
Resolution is the biggest lever once you leave Infer. On Infer, Veo 3.1 Fast tops out at 1080p at the flat $0.09-$0.10/sec rate ; there's no 4K option to upgrade into on this platform. Off Infer, both Google's own API and fal.ai gate 4K behind a 2.5-3x per-second premium over 1080p.
Duration works differently than most video models here: Veo 3.1 Fast generates natively in 8-second segments rather than a continuous slider, and Infer documents chaining those segments up to roughly 148 seconds at the same per-second rate. No volume discount, no penalty, just the same $0.09-$0.10/sec repeated across however many 8-second units you stitch together.
Billing only applies to generations that complete. Infer's broader platform pricing states the API is "pay only for what you use," and the model catalog bills per second of successful output rather than per attempt, consistent with how Infer prices its other video models .
Cheaper alternatives if audio isn't the requirement
If native synchronized audio isn't the reason you're on Veo 3.1 Fast, Kling 3.0 1080p Pro sits at the same $0.10/sec on Infer, holds native 4K capture where Veo caps at 1080p, and outranks it on Infer's leaderboard (1248 vs. 1208 Elo, snapshot 2026-04-30). The tradeoff is total: Kling generates silent clips, full stop, so it only works if you don't need dialogue or Foley baked into the render. See the full Kling vs. Veo comparison for the test-by-test breakdown.
For volume work where audio still matters, Seedance 2.0 Pro runs $0.13/sec with joint audio-video generation and a higher Elo (1271) than either Veo or Kling, though its Infer listing caps at 720p rather than Veo's 1080p. Run Seedance 2.0 Pro on Infer →
Infer's take rate doesn't change based on which of these three you pick, including the two alternatives just recommended above, so pointing you at Kling or Seedance instead of the model this article is nominally about isn't a soft sell for anything.
Related reading
- The cheapest AI video generation APIs in 2026: where Veo 3.1 Fast lands against Hailuo, Kling, Wan, and Seedance on price alone.
- Kling 3.0 Pro vs Veo 3.1 Fast: the full quality-and-price verdict between these two.
- The AI video pricing index: every model and provider rate in one table.
- Best AI video models in 2026: the ranked list Veo 3.1 Fast sits on.
- All pricing and cost breakdowns: every calculator on the site.
Frequently asked questions
Does Veo 3.1 Fast cost extra for audio?
Not on Infer or on Google's own Gemini API. Infer's $0.09-$0.10/second rate includes audio with no separate line item, and Google's official pricing page lists a single 'video with audio price (default)' for Veo 3.1 Fast rather than splitting silent and audio rates. fal.ai is the exception: it charges $0.10/second without audio versus $0.15/second with audio at 720p/1080p, a $0.05/second surcharge that neither Infer nor Google's direct API applies.
Does Veo 3.1 Fast support 4K?
Not on Infer, where the model's listed ceiling is 1080p. Google's own API does offer a 4K tier for Veo 3.1 Fast at $0.30/second, 2.5 to 3 times the 720p/1080p rate of $0.10-$0.12/second.
How much does a full-length chained clip cost?
Veo 3.1 Fast generates natively in 8-second segments, chainable to roughly 148 seconds on Infer. At Infer's $0.09-$0.10/second rate, 148 seconds costs $13.32-$14.80. The same duration through Google's own API at the Fast 1080p rate ($0.12/second) would run $17.76.
Is Veo 3.1 Fast cheaper on Infer than going direct to Google?
Yes, at 1080p. Infer's $0.09-$0.10/second beats Google's own Fast-tier 1080p rate of $0.12/second, confirmed directly against Google's Gemini API pricing page. At 720p the two are close to parity ($0.09-$0.10 on Infer vs. $0.10 on Google's page).
Sources