|
English

The generative media market crossed a threshold this year that nobody announced. Producing a good-looking frame stopped being the hard part. Every serious lab can now make a still that survives a client review, and most can make several seconds of video that survives one too. What separates the models in August 2026 is not whether they can produce the shot. It is how many attempts you burn before you get one you can use, and what happens to your pipeline when a vendor retires the model you built it on.

That second question is not hypothetical right now, which is where this comparison has to start.

The thing to deal with first: Sora 2 is going away

OpenAI discontinued the Sora app on 26 April 2026, and the API has a confirmed shutdown date of 24 September 2026. That is weeks away, not quarters. Sora 2 remains accessible to ChatGPT Plus and Pro subscribers inside ChatGPT, but as an API dependency it is finished.

If you have a pipeline calling the Sora API, migration is now a scheduled piece of work rather than a risk to monitor. If you are evaluating video models this month, Sora 2 should not be on the shortlist regardless of how it scores, and any comparison that still ranks it as a live option is out of date.

This is worth dwelling on because it is the clearest recent argument for not committing a production pipeline to a single generator. The model that led the category eighteen months ago is being switched off, and Runway Gen-4.5, which led on Elo at launch in late 2025 with 1247, has since dropped out of the top ten entirely. Nothing about a leadership position in this category has proven durable.

Video models

Per-second rates at 1080p unless noted, as of early August 2026. Prices in this category re-quote frequently, so treat these as a starting point and check the provider's live page before committing a budget.

ModelPrice / secNative audioNotes
Sora 2 Pro~$0.75YesAPI shuts down 24 Sept 2026
Veo 3.1 (4K)$0.60Yes48kHz synchronised dialogue
Veo 3.1 (720p/1080p)$0.40YesFast mode drops to $0.15
Sora 2~$0.10YesAPI shuts down 24 Sept 2026
Kling 3.0~$0.07 to $0.10YesNative 4K, 60fps, 15s clips
Veo 3.1 Lite~$0.05YesLowest entry price
Wan 2.6~$0.05NoOpen source

On the leaderboards, ByteDance's Seedance 2.0 and Alibaba ATH's HappyHorse-1.0 hold the top two slots on Artificial Analysis. Veo 3.1 sits at number three and is the leader among models with reliable enterprise API access. Kling 3.0 has four separate entries inside the top ten, which is a better signal of a mature family than any single score.

Head-to-head results diverge in a way worth knowing about. On Elo, Kling 3.0 1080p Pro sits around 1,104 and edges out Veo 3.1. On a hands-on comparison published in June 2026, the gap was far wider than Elo suggests, with Kling's Pro tier scoring 91 out of 100 against Veo 3.1 at 63 and Sora 2 Pro at 59. Preference Elo and structured review disagree here, and when two methodologies disagree that sharply the honest conclusion is that the ranking depends on what you are making rather than that one model is better.

What each is actually good at: Veo 3.1 still owns synchronised dialogue, generating 48kHz speech aligned to the footage, which matters enormously if your output involves people talking and not at all otherwise. Kling 3.0 gives you native 4K at 60fps, fifteen-second clips and multilingual lip-sync, and it is four to seven times cheaper than the premium tier. Wan 2.6 is the open-source floor, giving up quality for control and a price nothing hosted can match.

Resolution has stopped being the interesting axis. Every serious model does 1080p or native 4K now. The axes that still separate them are audio, clip length, lip-sync quality and price.

Image models

Image pricing sits low enough that the spread rarely decides anything, but the range is wider than people assume.

ModelPriceStrongest at
Nano-Banana Pro~$0.20 / imagePremium generation plus editing
Ideogram v3 Balanced~$0.07 / imageTypographic design
Imagen 4~$0.05 / imageText rendering accuracy, speed
Seedream 4.5~$0.04 / imageComposition
Flux 2 Pro~$0.02 / image, ~$0.03 / MPPhotorealism, prompt adherence
FLUX.1 schnell~$0.003 / MPVolume, speed
GPT Image 2~$0.002 to $0.003Budget volume work

Flux 2 Pro is the sensible default to evaluate first, combining photorealism, prompt adherence and competitive pricing, and it accepts up to eight reference images which matters for consistency across a set. Imagen 4 leads on text rendering accuracy, which sounds niche until you try producing anything containing a logo, a sign, a label or a UI mock. Imagen 4 Ultra produces the most photorealistic output available. Ideogram v3 owns typographic design specifically. GPT Image 1.5 handles complex scene composition better than its rivals.

Two newer entries worth watching: Seedream 5.0 from ByteDance integrates real-time web search into the generation pipeline, which is a genuinely different capability rather than a quality increment, and Nano-Banana Pro bundles editing with generation at the top of the price range.

One billing detail that matters at volume. Replicate typically charges per second of GPU compute while fal.ai charges per image or per megapixel, so any comparison across those platforms needs normalising to cost per 1024x1024 image before it means anything. On fal, server errors in the 500 range are not billed, which is a small thing until you are generating at scale and hitting a bad hour.

The unit that matters: cost per usable shot

Per-second pricing is close to useless on its own, because it prices generation rather than results. The number to track is cost per usable shot.

Take a ten-second shot. Veo 3.1 at $0.40 per second costs $4.00 per attempt. Kling 3.0 at $0.10 costs $1.00. Suppose Veo lands a usable shot every 2.4 attempts on ordinary commercial work and Kling needs 3.5.

Veo 3.1: 2.4 x $4.00 = $9.60 per usable shot Kling 3.0: 3.5 x $1.00 = $3.50 per usable shot

Kling wins by nearly three times despite needing more tries. Now run the same arithmetic on a shot requiring unusual motion, where the premium model holds its hit rate at 3 attempts and the cheaper one climbs to 9:

Veo 3.1: 3 x $4.00 = $12.00 Kling 3.0: 9 x $1.00 = $9.00

The advantage nearly disappears, and once you add artist time spent reviewing nine clips instead of three, the cheap model has become the expensive one. A creative director at $85 an hour spending ten extra minutes per shot adds about $14 in labour, which exceeds both generation figures.

Those hit rates are illustrative, not measured. Nobody can publish a number that predicts yours, because it depends on what you are making and how specific your art direction is. What the arithmetic shows is the shape of the answer: cheap models win decisively on ordinary work and lose on hard work, and the crossover point is a property of your shot mix rather than of the models.

Measuring your own hit rate

You can establish yours in about a fortnight of ordinary work, and once you have it every purchasing argument becomes arithmetic instead of opinion.

The method is unglamorous. Tag every generation with a shot category before you send it, using categories that reflect your work rather than the vendor's demo reel. Something like talking head, product on white, environment establishing, character action, abstract motion. Five to eight buckets is plenty. Log the model, the attempt number, and a single binary from whoever reviews it, which is whether the output went into the edit or did not.

After two weeks, divide attempts by acceptances per category per model, multiply by the rate, and you have cost per usable shot broken down by the kind of work you actually do. In every team we have run this with, the result contained at least one surprise, and the surprise was rarely the model anyone expected to win.

The log also gives you a prompt-quality signal worth as much as the model comparison. If one operator's hit rate is double another's on the same category and the same model, the difference is craft rather than tooling, and it is transferable. Generative media has developed the same dynamic that routing developed for language models, where operational discipline around the model produces larger savings than swapping the model does.

Keep the log running afterwards. Hit rates move when vendors ship updates, usually without an announcement.

Building the evaluation set

One more piece of discipline, because the shortlist you start from determines everything downstream.

Most teams evaluate video models on the prompts vendors put in their demo reels: dramatic camera moves, striking lighting, unusual subjects. Those prompts are chosen because they show a model at its best, and they have almost no overlap with the shots a business actually needs. Evaluating on them tells you which model makes the best showreel, which is a question nobody is paying you to answer.

Build the set from your own back catalogue instead. Take thirty shots you have delivered in the last year, write the prompt you would need to regenerate each, and run every candidate against all thirty. It is a day of work and it replaces months of arguing from impressions.

Include the boring ones deliberately. A product rotating on white, a person talking to camera, an establishing exterior. These are where most volume sits and where models differ least in demos and most in practice, particularly on the details that force a reshoot: hands, text on packaging, brand colours drifting between takes, and whether a face stays the same face across a cut.

Score on a single binary rather than a rubric. Would this go in the edit, yes or no. Multi-point aesthetic scales feel more rigorous and produce worse decisions, because they let a model that never quite clears the bar accumulate a respectable average.

Licensing, provenance and the question your lawyer will ask

Every model here will happily generate you something. Fewer will tell you clearly what you may do with it.

Three questions to put to any vendor before committing a campaign. Who owns the generated asset, and does that ownership survive if you fine-tuned on your own material? What indemnity, if any, covers you if a rights holder claims the output resembles their work? And does the model embed C2PA content credentials, since a growing number of platforms and broadcasters now require them and provenance cannot be retrofitted after generation.

Google and OpenAI both offer commercial indemnity on enterprise tiers. Indemnity positions elsewhere are thinner, and for regulated advertising work that difference is worth more than a price advantage. Open-weight options give you no indemnity at all, which is the honest trade for the control they hand you.

The other consideration is where generation happens. If unreleased product imagery is going through a third-party API, it has left your perimeter, and for some industries that ends the conversation before quality is discussed. This is the same data sovereignty calculus that governs text workloads, and it is why Wan 2.6 keeps appearing on shortlists despite being a step behind.

What to actually run

For a small team producing social and marketing assets, Flux 2 Pro for stills and Kling 3.0 for video, with Veo 3.1 available for shots that need dialogue. Expect that to cover you for well under half of what a single premium vendor would cost.

For an agency with client indemnity requirements, Imagen 4 and Veo 3.1 on enterprise terms, with indemnity as the deciding factor rather than the quality delta.

For anyone with confidential source material or a self-host mandate, Wan 2.6 and an open image model, accepting that you are a generation behind and that you own the serving.

Whichever combination you land on, put it behind one interface rather than four sets of credentials. Aggregators like fal.ai and Replicate exist for this, and Swfte Connect treats image and video endpoints the way it treats language models, with per-request routing, one bill, and the attempt logging you need to compute cost per usable shot rather than guessing at it.

Sora 2's shutdown is the argument in a sentence. The models will keep changing. The discipline that survives is measuring your own hit rate on your own work, because that number is specific to you and no benchmark will ever tell you what it is.


Related: Best AI video generators ranked by quality, the generative media wave moving into geometry, and the live model leaderboard.

0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.