The State of AI Video Generation: July 2026

By the Infer teamUpdated

Video generation Elo scores moved by triple digits, audio went from a Veo-only party trick to a five-model standard, and the single most-discussed model of the year still has no API. This report compiles what the data actually shows across Infer's model catalog, Artificial Analysis's arena, and provider pricing pages as of July 2026: cite it, argue with it, or use it to sanity-check the next hot take about "the best AI video model."

Key findings

  • The top-ranked video model in the world has no public API. HappyHorse-1.0 sits #1 on Infer's cached Artificial Analysis snapshot at 1368 Elo, more than 90 points clear of #2, and Infer's own leaderboard page notes it as an "anonymous top arena entry; Alibaba-confirmed authorship, no public API yet" (tryinfer.com/leaderboards, snapshot dated 2026-04-30).
  • Leaderboards are snapshots, not truths. Kling 3.0 1080p Pro ranks #3 out of 12 tracked models at 1248 Elo on the 2026-04-30 Artificial Analysis data Infer cites, and #6 at 1111 Elo on Artificial Analysis's own live arena as of 2026-07-22: the same model, an 83-day gap, a three-rank and 137-point swing (artificialanalysis.ai/video/leaderboard/text-to-video).
  • Native audio-video generation is now table stakes, not a differentiator. Veo 3 shipped the first mainstream native-audio model on May 20, 2025; about eight months later, Seedance 2.0, Runway Gen-4.5 (December 1, 2025), Wan 2.6 (December 16, 2025), and the open-weight LTX-2 (January 6, 2026) all generate synchronized audio and video in a single pass (runwayml.com/research/introducing-runway-gen-4.5, alibabacloud.com Wan2.6 press room).
  • China builds the volume leader. Alibaba (HappyHorse, Wan), ByteDance (Seedance), and Kuaishou (Kling) models fill 4 of the top 6 slots on Infer's cached leaderboard and 7 of the top 10 on Artificial Analysis's live arena as of 2026-07-22.
  • Price per second of video spans roughly 8x across quality tiers. From $0.08/sec for Hailuo 02 Pro on Infer to $0.682/sec for Seedance 2.0 at 1080p standard tier on fal.ai, the same underlying task, an 8.5x price gap depending on host and tier (tryinfer.com/models/hailuo-02-pro, fal.ai/seedance-2.0).
  • Open weights run one generation behind closed models on quality, but LTX-2 closed the feature gap first. Wan 2.2 and LTX trail Kling 3.0 and Seedance 2.0 on Elo, yet LTX-2, released January 6, 2026, became the first production open-weight model with native synchronized audio+video generation, ahead of several closed competitors adding audio at all (globenewswire.com LTX-2 release).

Methodology

This report draws on three live data sources, each dated independently rather than blended into one number:

  1. Infer's leaderboard page (tryinfer.com/leaderboards), observed 2026-07-22, which itself cites Artificial Analysis Video Arena Elo data as of 2026-04-30 and VBench Leaderboard percentages as of 2026-04-17.
  2. Artificial Analysis's live text-to-video arena (artificialanalysis.ai/video/leaderboard/text-to-video), fetched directly on 2026-07-22: a different, more current snapshot than the one Infer's page cites.
  3. Provider pricing pages for the 14 models in Infer's catalog, plus fal.ai and official-API listings for cross-provider comparison, all observed 2026-07-22.

We deliberately do not merge the two leaderboard snapshots into a single ranked list; see finding two above for why that would misrepresent both. Every chart's underlying numbers appear as a markdown table directly beneath it, so you can check our arithmetic. Infer hosts all of these models and takes no side in how they rank against each other; this report reflects the data as fetched, not a promotional lean toward Infer's own catalog.

Where the leaderboards stand today

Infer's cached Artificial Analysis snapshot (Elo, as of 2026-04-30)

1HappyHorse-1.0Alibaba1368No public API
2Dreamina Seedance 2.0 (720p)ByteDance1271Native audio + phoneme-level lip-sync
3Kling 3.0 1080p (Pro)Kuaishou1248Native 4K, strong multi-shot consistency
4SkyReels V4Skywork AI1239
5Kling 3.0 Omni 1080p (Pro)Kuaishou1233
6grok-imagine-videoxAI1231
7Vidu Q3 ProVidu1221
8Runway Gen-4.5Runway1216
9Veo 3.1Google DeepMind1208
10Wan 2.6Alibaba1189
11LTX-2 ProLightricks1129
12LTX-2.3 FastLightricks1121

Source: tryinfer.com/leaderboards, observed 2026-07-22, citing Artificial Analysis data dated 2026-04-30.

Artificial Analysis live arena (Elo, observed 2026-07-22)

1Gemini Omni FlashGoogle1243
2Dreamina Seedance 2.0 (720p)ByteDance Seed1227
3Wan2.7-260612Alibaba1164
4HappyHorse-1.1Alibaba-ATH1151
5HappyHorse-1.0Alibaba-ATH1128
6Kling 3.0 1080p (Pro)KlingAI1111
7SkyReels V4Skywork AI1108
8Wan 2.7Alibaba1107
9Kling 3.0 720p (Standard)KlingAI1100
10Veo 3.1Google1096

Source: artificialanalysis.ai/video/leaderboard/text-to-video, observed 2026-07-22.

Put the two tables side by side and the takeaway isn't "which one is right": both are, for the day they were pulled. Seedance 2.0 (720p) is the one model that holds a top-2 spot on both, which is the closest thing to a stable signal in this data. Everything else, including HappyHorse's exact rank, whether Kling sits third or sixth, and whether Gemini Omni Flash or grok-imagine-video even appear, depends on which week you looked. Six models on the live AA arena (Gemini Omni Flash, Wan2.7-260612, HappyHorse-1.1, Wan 2.7, Kling 3.0 720p Standard, and the reshuffled Veo position) don't match the composition of Infer's cached table at all. If you're citing a leaderboard rank in a pitch deck three months from now, it's already wrong. Go re-pull it.

Audio went from party trick to requirement

Veo 3May 20, 2025Google DeepMind launch coverage
Runway Gen-4.5December 1, 2025runwayml.com research post
Wan 2.6December 16, 2025Alibaba Cloud press room
Seedance 2.0Live on fal.ai by ~April 2026fal.ai/seedance-2.0
LTX-2 (open weights)January 6, 2026globenewswire.com LTX-2 release

About eight months separate Veo 3's May 2025 launch, the first mainstream model to generate audio and video from one prompt, from LTX-2's open-source release in January 2026, which brought the same capability to self-hostable weights. In between, Runway framed Gen-4.5's "native dialogue, ambient sound, and synchronized music as unified multimodal output" as a research breakthrough; about eight months later it's a checkbox every serious video model needs. Infer's own changelog logged "native audio toggling added across multiple models" on 2026-07-13, the same entry that rolled 1080p out to Seedance 2.0 Pro, Veo 3.1 Fast, HappyHorse 1.1, and WAN 2.2 Flash simultaneously (tryinfer.com/changelog). That's the industry treating audio and resolution as a joint feature release now, not separate milestones.

The holdout is Kling 3.0 1080p Pro, which still ships no native audio in the version tracked on both leaderboards. Kuaishou's own docs reference audio credit add-ons rather than joint generation, worth flagging if you're deciding between Kling and Seedance for anything with dialogue.

China ships the volume

Four of the top six models on Infer's cached snapshot come from Chinese developers: HappyHorse-1.0 (Alibaba, #1), Dreamina Seedance 2.0 (ByteDance, #2), Kling 3.0 1080p Pro and Kling 3.0 Omni (Kuaishou, #3 and #5). On Artificial Analysis's live arena the pattern holds even tighter: seven of the top ten, Seedance 2.0, Wan2.7-260612, both HappyHorse entries, Wan 2.7, and both Kling variants, trace back to Alibaba, ByteDance, or Kuaishou. Google and xAI hold the remaining top-10 slots between the two tables; no US or European lab currently places more than one model in either top ten. This isn't a story about any single lab's algorithm. It's four Chinese developers iterating on multiple concurrent model families (Kling has three variants on one table alone) while most Western labs field a single flagship at a time.

What a second of video actually costs

Hailuo 02 ProInfer$0.08/sectryinfer.com/models/hailuo-02-pro
Veo 3.1 FastInfer$0.09–$0.10/sectryinfer.com/models/veo-3-1-fast
Kling 3.0 1080p ProInfer$0.10/sectryinfer.com/models/kling-3-0-pro
Wan 2.2 T2V-A14BInfer$0.13/sectryinfer.com/models/wan-2-2-t2v-a14b
Seedance 2.0 ProInfer$0.13/sectryinfer.com/models/seedance-2-0-pro
Seedance 2.0 (fast tier, reference-to-video)fal.ai$0.2419/secfal.ai/seedance-2.0
Seedance 2.0 (720p standard)fal.ai$0.3034/secfal.ai/seedance-2.0
Seedance 2.0 (1080p standard)fal.ai$0.682/secfal.ai/seedance-2.0

The floor and ceiling of this table are the same underlying capability class, full text-to-video generation, separated by 8.5x. Some of that gap is real quality difference (Seedance's 1080p standard tier on fal.ai is a different render path than Infer's $0.13 rate for the same model), and some of it is just markup: hosting the identical model can cost 2–5x more depending on which API sits in front of it. Infer's rates in this table are the low end across every row we could verify directly; treat the fal.ai figures as the "official reseller" reference point for what unmanaged access costs. We flag but don't include several competitor prices from this comparison (official Google and Kuaishou API rates for Veo and Kling) because they reached us only through search-aggregator summaries rather than a primary pricing page. Re-verify those specifically before repeating them as fact.

Open weights: closing the feature gap, not the quality gap

Wan 2.2 T2V-A14B, the open-weights model Infer hosts, doesn't appear on either Elo leaderboard above; its developer positions it as "~85% quality compared to Hailuo 02 Pro" by Infer's own comparative copy, not as a top-12 arena contender. LTX-2 Pro and LTX-2.3 Fast do appear, at #11 and #12 on Infer's cached snapshot (1129 and 1121 Elo), solidly behind Seedance, Kling, and Veo, roughly a full model generation on Elo terms.

But LTX-2's real claim isn't its Elo rank. It's the date. Lightricks open-sourced it on January 6, 2026 as, in the company's own words, "the first production-ready model with truly open audio+video weights": full model weights, inference pipeline, and training code, running native 4K at 50fps with synchronized audio in a single diffusion pass. Alibaba's closed Wan 2.6 beat it to audio-visual sync by three weeks (December 16, 2025), but LTX-2 was the first to ship that capability with the weights themselves public, and LTX-2.3, a 22-billion-parameter refinement, followed on March 5, 2026 with sharper output and better portrait support. If you're building on open weights and need audio, LTX-2 is currently the only entry in that category with a shipping date instead of a roadmap slide.

What changes next

  • The leaderboard gap between Infer's cached snapshot and Artificial Analysis's live arena will keep widening, then reset. Expect any single "top model" claim published today to be stale within 4–8 weeks based on the 83-day, 137-point swing already observed for Kling 3.0 1080p Pro between April and July 2026.
  • Audio-native generation becomes a baseline requirement for any new flagship launch by early 2027. Every major release since Veo 3 (May 2025) has shipped it; a text-only launch from a top-tier lab would now read as a regression, not a feature choice.
  • Open-weight models will keep trailing frontier Elo by roughly one generation, but the feature-parity gap will keep shrinking. LTX-2 hit synchronized audio+video eight months after Veo 3 did; expect the next open-weight capability (extended duration, higher native resolution) to arrive within a similarly compressed window rather than the multi-year lag open models showed in 2023–2024.
  • Price compression will hit the $0.13/sec tier hardest. Seedance 2.0 Pro and Wan 2.2 currently anchor the top of Infer's price ladder at $0.13/sec; with fal.ai's own Seedance rates already ranging up to $0.682/sec for a comparable render, there's more room for a hosted price war at the high-quality tier than at the already-compressed $0.08–$0.10 value tier.

Cite this report

Citation: Infer, "The State of AI Video Generation: July 2026," tryinfer.com, published July 28, 2026. Data compiled from Infer's leaderboard (2026-04-30 Artificial Analysis snapshot), Artificial Analysis's live text-to-video arena (observed 2026-07-22), and provider pricing pages (observed 2026-07-22). Canonical URL: https://tryinfer.com/guides/state-of-ai-video-2026

Run any model in this report, from Hailuo 02 Pro at $0.08/sec to Seedance 2.0 Pro at $0.13/sec, through Infer's playground, one API key for the full catalog.

For the deeper cuts behind these numbers, see the cheapest AI video APIs, ranked, the best AI video models in 2026, Kling 3.0 vs Veo 3.1, Seedance 2.0 vs Kling 3.0, Wan vs Kling vs Hailuo, and the full guides hub.

Frequently asked questions

What is the best AI video model right now, by Elo score?

HappyHorse-1.0 leads Infer's cached Artificial Analysis snapshot (2026-04-30) at 1368 Elo, but it has no public API. Among models you can actually call, Dreamina Seedance 2.0 (720p) is highest at 1271 Elo.

Why do AI video leaderboards disagree with each other?

Because they're snapshots, not fixed rankings. Kling 3.0 1080p Pro sits #3 at 1248 Elo on the 2026-04-30 Artificial Analysis data Infer's leaderboard cites, but #6 at 1111 Elo on Artificial Analysis's own live page as of 2026-07-22; new entrants and re-ratings reshuffle the board every few weeks.

Do all AI video models generate audio now?

Not all, but the frontier has moved there fast. Veo 3.1, Seedance 2.0, Runway Gen-4.5, Wan 2.6, and the open-weight LTX-2 all ship native audio-video generation as of July 2026, up from essentially zero mainstream models 12 months earlier.

Which country is producing the most AI video models?

China. Alibaba, ByteDance, and Kuaishou models occupy most of the top-10 slots on both Infer's cached leaderboard and Artificial Analysis's live arena as of July 2026.

How wide is the price spread across AI video providers, not just Infer?

Roughly 8x market-wide: $0.08/sec (Hailuo 02 Pro on Infer) to $0.682/sec (Seedance 2.0 1080p standard tier on fal.ai) across the providers and quality tiers in this report.

Are open-weight video models catching up to closed ones?

They're closing the gap on features faster than on quality. LTX-2, released January 6, 2026, became the first open-weight model with native audio-video generation, but the top open entries still trail Kling and Seedance on Elo.

Sources

Related reading