How AI video leaderboards actually work (and why they disagree)

By the Infer teamUpdated

An AI video leaderboard is a snapshot of blind human votes on a specific day, not a fixed ranking of model quality. That's the whole explanation for why Kling 3.0 1080p Pro sits #3 out of 12 models at 1248 Elo on the Artificial Analysis data Infer's own leaderboard page cites (dated 2026-04-30), and #6 out of 10 at 1111 Elo on Artificial Analysis's live arena page as of 2026-07-22. Same model, an 83-day gap, a three-rank and 137-point swing (artificialanalysis.ai/video/leaderboard/text-to-video). This page explains what Elo actually measures, how Artificial Analysis's Video Arena and VBench differ, why self-reported numbers never match third-party ones, and how Infer's own leaderboard and rankings actually use this data.

Key findings

  • The same model, Kling 3.0 1080p Pro, ranks #3 on one AA snapshot and #6 on another 83 days later: 1248 Elo (2026-04-30, cited by Infer's leaderboard) vs. 1111 Elo (2026-07-22, AA's live page). No error occurred; both numbers are correct for the day they were pulled (tryinfer.com/leaderboards, artificialanalysis.ai/video/leaderboard/text-to-video).
  • Elo ratings on Artificial Analysis recompute hourly, so any published rank is, at best, a few hours old the moment it's screenshotted .
  • The rating math is Bradley-Terry, rescaled to look like Elo, not the original chess formula .
  • Self-reported scores run higher than third-party ones for the same model. Runway states Gen-4.5 achieved 1247 Elo, "ranking first on global text-to-video leaderboards," per its own research post, while Infer's cached third-party (Artificial Analysis) snapshot has Runway Gen-4.5 at #8, 1216 Elo: a 31-point gap between the lab's number and the arena's.
  • VBench scores on 16 separate dimensions instead of one blended number. Subject consistency, motion smoothness, and imaging quality are graded independently rather than folded into a single win/loss vote .
  • Three models on Artificial Analysis's live arena don't appear anywhere on the snapshot Infer's leaderboard cites: Gemini Omni Flash (now #1, 1243 Elo), Wan2.7-260612 (#3, 1164 Elo), and HappyHorse-1.1 (#4, 1151 Elo) were all added after the 2026-04-30 cutoff.
  • The top-ranked model on Infer's cached snapshot, HappyHorse-1.0 at 1368 Elo, has no public API. It's an anonymous arena entry Infer's own page attributes to Alibaba, with no way for outside testers to independently verify the result.

Methodology: what this page documents, and how we compiled it

This isn't a ranking page. It's the reference page every other Infer article links to when it cites a leaderboard number. Three sources feed it: Artificial Analysis's published Video Arena methodology (fetched 2026-07-22), Artificial Analysis's live text-to-video leaderboard compared against Infer's own cached leaderboard page, and VBench's public documentation for the dimension-scored alternative. Where a claim comes from Infer's own data pack, it's cited as such; where it comes from a lab's own research post rather than a third-party measurement, that distinction is called out explicitly, because that distinction is most of this article's point. Infer hosts every model discussed here and takes no side in how they rank against each other; this page explains the scoreboard, not who's winning on it.

What Elo actually is

Elo is a relative rating system invented for chess in the 1960s: two players face off, the winner takes points from the loser, and beating a higher-rated opponent earns more than beating a weaker one. There's no absolute "quality" score, only a number that moves based on who beat whom. Applied to AI video, Artificial Analysis's Video Arena replaces chess matches with blind pairwise comparisons: two models generate video from the same prompt, a human votes for the one they prefer without knowing which model made which clip, and that single vote nudges both models' ratings up or down .

The underlying math isn't the original 1960s Elo formula. It's Bradley-Terry Maximum Likelihood Estimation, a related but distinct statistical model for pairwise comparison data, rescaled afterward to produce numbers that look and behave like Elo ratings . The practical upshot for readers: a 137-point Elo gap doesn't mean "37% better" or any fixed percentage. It means the higher-rated model would be expected to win a majority of blind head-to-head votes against the lower one, at whatever margin the underlying win rate implies.

Under the standard Elo convention (expected win rate = 1 / (1 + 10^(−gap/400))), the translation looks like this — worth internalizing before you treat any two adjacent leaderboard rows as meaningfully different:

20 points~53%
40 points~56%
63 points (Seedance 2.0 vs Veo 3.1, Infer's 2026-04-30 snapshot)~59%
100 points~64%
137 points~69%

Read that middle row again: the #2 model beats the #9 model in blind voting about six times out of ten. Four times out of ten, human voters preferred the "worse" model's clip. A 23-point gap, like Seedance vs Kling on the same snapshot, is a 53/47 split — statistically real across thousands of votes, and close to a coin flip on any single prompt you personally run. This is why we treat leaderboard position as one input alongside price and documented-spec analysis, not a verdict. (Caveat: Artificial Analysis rescales Bradley-Terry estimates to an Elo-like range, so treat these conversions as the standard convention's approximation, not AA's published math.)

One structural detail matters if you're comparing models across use cases: Artificial Analysis scores each modality separately. Text-to-video, image-to-video, video editing, and the audio variants of each are rated as independent arenas, because matchups only ever pair outputs from the same modality . A model's text-to-video Elo says nothing about its image-to-video Elo, so every table in this article and elsewhere on Infer specifies which modality it's quoting.

VBench: scoring 16 dimensions instead of one vote

VBench takes a different approach entirely. Rather than a single blind-vote Elo number, it breaks "video generation quality" into 16 separately scored dimensions: subject consistency, background consistency, temporal flickering, motion smoothness, dynamic degree, aesthetic quality, imaging quality, object class accuracy, multiple-object handling, human action fidelity, color, spatial relationship, scene, temporal style, appearance style, and overall consistency . Each dimension gets its own automated evaluation pipeline rather than a crowd of humans voting on an overall winner. That's the core methodological split from an arena: VBench measures specific, named properties objectively, while Artificial Analysis measures aggregate human preference.

The two aren't competing measurements of the same thing. They answer different questions. VBench can tell you a model scores well on motion smoothness but poorly on multiple-object handling, a diagnosis an Elo number can't give you. An arena Elo can tell you which model people actually preferred watching, which a per-dimension score can't, because a model can win every individual dimension and still produce clips that feel worse in aggregate than the sum of its parts suggests. Infer's own leaderboard cites VBench as a secondary source alongside Artificial Analysis specifically because the two catch different failure modes.

Self-reported vs. third-party: why a lab's own number is never the number to cite

Labs publish their own benchmark numbers in launch posts, and those numbers are almost always higher, or framed more favorably, than what an independent arena later measures. Runway's own research post states Gen-4.5 achieved 1247 Elo, describing it as "ranking first on global text-to-video leaderboards" at launch. Infer's leaderboard page, sourced from Artificial Analysis's third-party arena, places Runway Gen-4.5 at #8 overall with 1216 Elo: 31 points and seven ranks below the self-reported figure. Neither number is fabricated; they measure different vote pools at different times against different competitor sets, and a lab's launch-day claim is measured under conditions the lab itself controls (which competitors it benchmarks against, which prompt set, which checkpoint). Third-party arena figures are measured against whatever roster happens to be live on the arena that day, which is a fairer comparison but a moving target.

The practical rule this article follows, and the one worth adopting whenever you read a model announcement: treat a lab's own benchmark claim as a marketing floor, not a verified ceiling, and look for the third-party arena number before repeating a "#1" claim.

The centerpiece: Kling 3.0 Pro's two rankings

This is the clearest illustration of the whole problem, because it's the same exact model measured by the same benchmark provider, 83 days apart.

Infer's cached Artificial Analysis snapshot#3 of 121248tryinfer.com/leaderboards, dated 2026-04-30
Artificial Analysis live arena#6 of 101111artificialanalysis.ai/video/leaderboard/text-to-video, observed 2026-07-22

Nothing changed about Kling 3.0 1080p Pro's weights between April 30 and July 22 that we know of. What changed is the competitive field it's being measured against. Three models that didn't exist on the April snapshot now sit above it on the live arena: Gemini Omni Flash (#1, 1243 Elo), Wan2.7-260612 (#3, 1164 Elo), and HappyHorse-1.1 (#4, 1151 Elo). Kling didn't get worse. The field around it got more crowded, and Elo is a relative measure, so an unchanged model's rating drifts as new, well-rated entrants absorb points from every model they beat. Six of the ten models on the current live arena don't match the composition of the snapshot Infer's own page cites at all.

We deliberately never blend these two tables into one ranking. Doing so would imply a single, continuous "Kling's Elo over time" trend line, when what actually exists is two independent measurements with different rosters, different vote pools, and an 83-day gap. If you're citing either number, cite its date and its source page alongside it. "Kling 3.0 Pro, #3 at 1248 Elo (Infer's leaderboard, April 2026 snapshot)" is a defensible sentence; "Kling 3.0 Pro, #3 at 1248 Elo" without a date is already out of date by the time you publish it.

How Infer sources its own leaderboard

Infer's leaderboard page (tryinfer.com/leaderboards) tracks three category tabs: video generation (12 models), image generation (9 models), and vision-language models (8 models). It cites Artificial Analysis Video Arena Elo as its primary video source, dated 2026-04-30 on the page as observed, plus VBench Leaderboard percentages dated 2026-04-17 as a secondary reference. The page states rankings are "refreshed regularly" and describes the underlying market as moving "fast," language that undersells how literally true it is given the 83-day swing documented above.

One caveat worth flagging for anyone trying to reproduce our numbers: only the video-generation tab's table renders in the page's initial HTML. The image-generation and vision-language tabs load their row data client-side on tab click, so the image and VLM leaderboard positions cited elsewhere on Infer (GPT Image 1.5 at #2, Nano Banana 2 at #3, Seedream 4.0 at #6) were recovered from each model's individual page, not pulled directly from a captured leaderboard table for those two categories. Treat any "top 9 image models" or "top 8 VLM" claim as provisional until someone re-derives it from a live browser session.

How we actually use leaderboards in our rankings

Infer hosts every model on every leaderboard discussed here and profits the same amount regardless of which one wins, so a leaderboard position alone never decides a ranking on this site. Our /best/ and /compare/ pages combine three inputs: the Artificial Analysis and VBench Elo figures documented above, each model's price on Infer, and documented-spec analysis: cross-checking each model's own product page and changelog for capabilities a leaderboard number can't capture. A model can lead on Elo and still lose a head-to-head recommendation: Kling 3.0 1080p Pro outranks Veo 3.1 Fast on both AA snapshots, but we still point readers to Veo when the job needs native synchronized audio, because no Elo number tells you which model has that capability at all.

This is also why HappyHorse-1.0's #1 position (1368 Elo, Infer's cached snapshot) doesn't get treated as "the best video model" anywhere else on this site: it has no public API, so nobody outside its arena votes can verify the result independently, and a ranking with no documented spec sheet or pricing page behind it doesn't earn a recommendation, whatever its Elo says.

Predictions: how this changes over the next year

  • Any leaderboard rank published today will be measurably stale within 8–12 weeks. The 83-day, 137-point swing on Kling 3.0 1080p Pro is the baseline case, not an outlier; expect similar drift on any model as new entrants join the arena.
  • Self-reported vs. third-party gaps will keep showing up at every major launch, because labs will keep benchmarking against whatever competitor set flatters a new release, while third-party arenas measure against the live field. Expect the gap to widen, not close, as more models launch with their own "#1" claims.
  • VBench-style dimension scoring will matter more as models converge on similar Elo bands. Once several models cluster within 10–20 Elo points of each other, which dimension a model is strong or weak on becomes more actionable than its aggregate rank. Expect more coverage that reports per-dimension scores alongside Elo, not instead of it.

Cite this report

Citation: Infer, "How AI Video Leaderboards Actually Work (and Why They Disagree)," tryinfer.com, published August 7, 2026. Data compiled from Artificial Analysis's Video Arena methodology page and live text-to-video leaderboard (both observed 2026-07-22), Infer's own leaderboard page (observed 2026-07-22, citing an Artificial Analysis snapshot dated 2026-04-30), and VBench's public documentation. Canonical URL: https://tryinfer.com/guides/how-ai-video-leaderboards-work

Test any model discussed here yourself rather than trusting a single Elo number. Kling 3.0 1080p Pro and Seedance 2.0 Pro are both live on Infer's playground with one API key.

For the numbers this methodology feeds into, see the State of AI Video Generation report, the best AI video models in 2026, Kling 3.0 vs Veo 3.1, Seedance 2.0 vs Kling 3.0, Wan vs Kling vs Hailuo, and the full guides hub.

Frequently asked questions

What is Elo, and what does it mean for an AI video model?

Elo is a rating system, borrowed from chess, that updates a score after each head-to-head result: beat a higher-rated opponent and you gain more points than beating a lower-rated one. Artificial Analysis applies it to AI video by having people vote blind on which of two model outputs they prefer for the same prompt, then converting the aggregate of those votes into an Elo-like number (artificialanalysis.ai/video/methodology).

Why do AI video leaderboards disagree with each other?

Because each one is a snapshot with its own cutoff date, vote pool, and model roster, not a single evolving truth. Kling 3.0 1080p Pro sits #3 at 1248 Elo on the 2026-04-30 Artificial Analysis data Infer's leaderboard cites, and #6 at 1111 Elo on Artificial Analysis's own live page as of 2026-07-22: same model, 83 days apart, a three-rank and 137-point swing.

Can AI labs game arena leaderboards?

Not the vote itself, since matchups are blind, but labs choose which checkpoint, prompt set, and modality tab to submit, and timing a submission right before a leaderboard snapshot is captured is a known soft-gaming tactic. HappyHorse-1.0 topped Infer's cached snapshot at 1368 Elo as an anonymous entry with no public API, a result nobody can independently stress-test yet.

How often do AI video leaderboard rankings change?

Artificial Analysis recomputes its Elo ratings hourly as new votes come in (artificialanalysis.ai/video/methodology), and the roster of models being voted on changes every few weeks: three models on the live arena (Gemini Omni Flash, Wan2.7-260612, HappyHorse-1.1) don't appear at all on the snapshot Infer's own leaderboard page cites from 2026-04-30.

Does Infer just rank models by leaderboard position?

No. Infer's rankings and recommendations combine leaderboard Elo (Artificial Analysis and VBench) with pricing and documented-spec analysis: what each model's own docs and changelog confirm it can and can't do. A model can lead on Elo and still lose a head-to-head recommendation on cost, reliability, or a documented capability gap, like missing native audio.

Sources

Related reading