How to read this

Artificial Analysis's Video Arena works like a chess rating: people are shown two AI-generated clips side by side, pick which one they prefer, and each model's Elo score moves up or down based on the outcome, exactly like a chess player's rating moves after a game. Higher Elo means more people preferred that model's output in head-to-head comparisons — it's a measure of human preference, not a technical spec like resolution or generation speed. Text-to-video and image-to-video are tracked as separate leaderboards, since a model can be much better at one than the other.

Text-to-video leaderboard

#ModelMakerElo
1Wan 3.0Alibaba1,238
2Gemini Omni FlashGoogle1,238
3Minimax H3 MaxMiniMax (via fal)1,235
4MiniMax H3MiniMax1,227
5Dreamina Seedance 2.0 720pByteDance Seed1,221
6Wan2.7-260612Alibaba1,157
7HappyHorse-1.1Alibaba-ATH1,146
10Kling 3.0 1080p (Pro)KlingAI1,108
Veo 3.1 & Grok ImagineGoogle / xAIoutside top 10

Image-to-video leaderboard

#ModelMakerElo
1Minimax H3 MaxMiniMax (via fal)1,201
2Dreamina Seedance 2.0 720pByteDance Seed1,190
3MiniMax H3MiniMax1,186
4Gemini Omni FlashGoogle1,180
5Wan 3.0Alibaba1,175
6Grok Imagine 1.5xAI1,109
7HappyHorse-1.1Alibaba-ATH1,105
Kling 3.0 1080p (Pro) & Veo 3.1KlingAI / Googleoutside top 10

Sourced directly from Artificial Analysis's Video Arena leaderboards (artificialanalysis.ai/video), re-verified September 4, 2026. Full leaderboards track 28+ models each — trimmed here to the ones relevant to tools covered on this site, plus the overall top 5-7 for context. These boards moved meaningfully in the three weeks since our last check (Wan 3.0 didn't exist on it a month ago) — treat every number here as a dated snapshot, not a permanent ranking.

What actually stands out

See Pricing for These Tools

Related: More AI video news · Wan 3.0 launch coverage · How AI videos are made · Why Sora isn't on this list

FAQ

What is the #1 AI video model right now?

As of September 2026, Alibaba's Wan 3.0 — released August 24, 2026 — has reached the top of Artificial Analysis's text-to-video leaderboard, running essentially tied with Google's Gemini Omni Flash, with Minimax H3 Max and MiniMax H3 close behind. On the separate image-to-video leaderboard, Minimax H3 Max currently leads, with Seedance 2.0 720p, MiniMax H3, Gemini Omni Flash, and Wan 3.0 all clustered within a few Elo points of each other. These rankings shift regularly as new models and versions release — Wan 3.0 itself wasn't on this board a month ago.

What is Wan 3.0, and why did it jump to #1?

Wan 3.0 is Alibaba's latest video-generation model, released August 24, 2026 through Alibaba Cloud's Model Studio and API, a day after Alibaba raised HK$80 billion (~$10.2 billion) in a Hong Kong share offering earmarked for AI infrastructure. It generates up to 30 seconds of video in a single pass (roughly double its predecessor, Wan2.7) from text, images, documents, slides, or web pages, and it debuted at or near the top of Artificial Analysis's text-to-video leaderboard almost immediately. See our dedicated coverage of the Wan 3.0 launch for pricing and what it's actually good at.

Where does Runway rank?

Runway's Gen-4.5 led Artificial Analysis's leaderboard at launch in late 2025 with a 1,247 Elo score, but had dropped out of the top 10 by May 2026 as newer models (Seedance 2.0, HappyHorse, the Kling 3.0 and Veo 3.1 cluster) overtook it. Runway's continued strength is its editing toolset layered on top of generation, not raw leaderboard position.

Is Grok Imagine actually the #1 AI video model, as some marketing claims?

No, not per Artificial Analysis's own published leaderboard. As of August 2026, Grok Imagine's video model ranks around the middle of the text-to-video board (#18 of 28 tracked models) and higher on image-to-video specifically for its newer 1.5 version (#4), but neither places it at #1 on either leaderboard.

How is this ranking measured?

Artificial Analysis's Video Arena uses head-to-head human preference voting between pairs of model outputs, converted into an Elo score — the same rating system used in chess. It's a measure of which output people preferred when shown two options side by side, not a technical benchmark of resolution, speed, or any single metric.