The best AI video generators, and where they still fall over

Every current model generates in four-to-ten-second pieces, Google publishes per-second prices you can do arithmetic with, and Coca-Cola needed 70,000 generated clips to assemble one commercial.

8 minute read

The biggest name in the category just switched off

Start with the most useful fact about generative video, because it is the one nobody leads with. OpenAI's own deprecations page records that on 24 March 2026 it notified developers that the Videos API and every sora-2 model alias would be removed on 24 September 2026, and the consumer Sora app was wound down ahead of that. TechCrunch, reporting the shutdown with app data from Appfigures, put lifetime in-app purchase revenue at roughly $2.1 million against a peak of 3.33 million downloads in November 2025. The model was not the problem; the product economics were. Treat that as the standing risk here. Build a workflow around a single generator and you are one company decision away from rebuilding it.

What is actually current, as of now

Google's Veo 3.1 is the model most creators meet first, because it sits inside the Gemini app and Flow as well as the API, and it generates audio natively rather than leaving you to score the clip. Runway's Gen-4.5 is the one for people who want controls rather than a prompt box: keyframes, image-to-video, video-to-video. Kling, MiniMax, ByteDance's Seedance and Alibaba's Wan family are all live and competitive, and Wan and Tencent's HunyuanVideo ship open weights you can run on your own hardware. Artificial Analysis runs a blind-vote arena where people compare two clips from the same prompt without knowing the source; in its mid-2026 text-to-video snapshot the top places go to Google, MiniMax, ByteDance and Alibaba, with Kling close behind. No model wins everything.

Everything still generates in four-to-ten-second pieces

Duration is the specification that shapes every other decision, and it has barely moved. Google's Veo documentation lists 4, 6 or 8 seconds per generation, with 1080p and 4K available only at the 8-second length and output in either 16:9 or 9:16. Runway's Gen-4.5 lets you pick anywhere from 2 to 10 seconds. Extension features exist but come fenced: Veo's extension outputs at 720p, and generated files are kept on Google's servers for only two days, so anything you intend to build on has to be extended inside that window. The practical consequence is worth saying plainly. Every AI video you have seen that runs longer than about ten seconds is an edit, not a generation. Someone made a pile of shots and cut them together.

What a second of generated video costs

Google publishes per-second prices, which makes this the one corner of the category where you can do honest arithmetic. On the Gemini API, Veo 3.1 Lite is $0.05 per second at 720p, Fast is $0.10, and standard Veo 3.1 is $0.40 per second at 720p or 1080p, rising to $0.60 at 4K. Gemini Omni Flash, Google's faster video model, works out at roughly $0.10 per second of 720p on the same page. Runway sells credits instead of seconds: its published plans are $12, $28 and $76 a month billed annually, carrying 625, 2,250 and 9,500 credits. The number that decides your budget is not the price per second, though. It is the price per second you actually keep.

Veo 3.1, 4K0.60 USD/secVeo 3.1, 720p/1080p0.40 USD/secVeo 3.1 Fast0.10 USD/secGemini Omni Flash0.10 USD/secVeo 3.1 Lite, 720p0.05 USD/sec
List prices from Google's published Gemini API pricing page, so every bar is the same vendor, same unit, same day. This is the price per second generated, which is not the number that decides a budget: at Coca-Cola's ratio of 70,000 clips for one finished ad, the cost that matters is the price per second you actually keep.

The ratio nobody quotes: 70,000 clips for one ad

TheWrap reported the production numbers behind Coca-Cola's 2025 AI-generated Christmas ad, taken from the brand's own making-of video. A team of five AI specialists worked through more than 70,000 generated clips across 30 days to assemble one commercial, and Coca-Cola's VP of generative AI said the craftsmanship was ten times better than the previous year's attempt. Whatever you make of the result, and the public reaction was rough two years running, that ratio is the honest picture of working with these tools at any scale. You do not prompt a shot. You fish for one. When you budget time or money for generated footage, budget for the discard pile rather than for the takes that survive.

Physics is still where it breaks

The clearest evidence that this is not a prompt-writing problem comes from Physics-IQ, a benchmark built by researchers at INSAIT and Google DeepMind. They filmed real physical events across solid mechanics, fluid dynamics, optics, thermodynamics and magnetism, then asked models to continue each one from its opening seconds. Testing Sora, Runway, Pika, Lumiere, Stable Video Diffusion and VideoPoet, they found physical understanding severely limited and, more usefully, unrelated to how realistic the output looked. A 2026 follow-up called Physics-IQ Verified rebuilt the benchmark's prompts and re-ran it on newer models including Sora 2, Wan 2.2 and HunyuanVideo 1.5; all of them still scored far below the ceiling you get by simply filming the same event a second time. Convincing and correct are independent variables.

Text, hands, and the continuity problem between shots

Three other failure modes come up constantly, and it is worth being straight that published measurement for them is thinner than for physics. On-screen text — signage, packaging, interfaces, anything with letterforms — garbles often enough that you should plan to fix it in post or frame it out. Hands doing something specific, particularly handling an object, stay unreliable. Continuity between shots is structural rather than cosmetic: each generation is an independent roll, so a jacket, a room's lighting and the colour of a mug drift between takes unless you pin them down. That every serious tool now ships reference images, ingredient images and start-and-end frame controls is the best available evidence of how real that problem is.

Where it earns a place in a creator's workflow

Two things to settle before you publish. YouTube requires creators to disclose realistic altered or synthetic content at upload and will apply the label itself if you do not, though clearly unrealistic or animated material and ordinary production assistance are exempt. And audiences are not neutral: the Coca-Cola backlash is what a brand cost looks like when the generated look is the point rather than the seasoning. For most creators the sane use is narrow — a few seconds of establishing or illustrative footage inside work that is mostly you, on camera, saying something worth hearing. That footage still has to be cut, reframed and captioned with everything else, which is the part FrameOS automates. 300 credits for 3 days · no card.

FAQ

What is the best AI video generator in 2026?

There is no single winner. Google's Veo 3.1 is the easiest to reach, sitting in the Gemini app, Flow and the API with native audio. Runway's Gen-4.5 gives you the most editing control. Kling, MiniMax, Seedance and Alibaba's Wan all rank near the top of Artificial Analysis's blind-vote arena, and Wan and HunyuanVideo ship open weights you can run yourself.

How long can AI-generated videos be?

Shorter than most people expect. Google's Veo 3.1 generates 4, 6 or 8 seconds per request, with 1080p and 4K limited to the 8-second length; Runway's Gen-4.5 covers 2 to 10 seconds. Longer pieces are assembled from many generations in an editor, with extension features that carry their own limits, so plan for editing rather than a single prompt.

How much does AI video generation cost?

Google's Gemini API lists Veo 3.1 Lite at $0.05 per second of 720p, Fast at $0.10, and standard Veo 3.1 at $0.40 per second at 720p or 1080p. Runway sells credits, with published plans at $12, $28 and $76 a month billed annually. Budget for the clips you throw away: Coca-Cola's 2025 AI ad reportedly went through more than 70,000 generated clips.

What can AI video generators still not do?

Physics, mainly. The Physics-IQ benchmark from INSAIT and Google DeepMind found physical understanding severely limited across every model tested, and unrelated to visual realism. On-screen text and hands handling objects remain unreliable, and continuity across shots has to be forced with reference images or keyframes because each generation is independent of the last.

Sources

Related pages