THE GENERATIVE MODEL MATRIX: YOUR LEADERBOARD WON’T FINISH THE FILM

THE GENERATIVE MODEL MATRIX: YOUR LEADERBOARD WON’T FINISH THE FILM

The September 2026 STAGES × NAKID comparison scores 43 models across video, image, 3D and realtime against what a production needs, not what a launch trailer promises.

A model can win an arena and still change your lead’s jacket between shots. It can generate an immaculate product image and misspell the only word the client actually cares about. It can produce a gorgeous sculpture that becomes someone else’s problem the moment you ask it to move.

Congratulations on the leaderboard placement. We still need a usable asset.

This is the gap the STAGES × NAKID Generative Model Comparison Matrix exists to address. Public preference rankings can tell you which outputs people favor under particular conditions. They cannot tell you everything about getting a film, fashion campaign or commercial out the door. That requires looking beyond the image that wins the comparison and into the workflow that survives the revisions.

The distinction matters because artists do not work inside launch announcements. They work inside constraints. A character needs to remain recognizable. An approved composition needs to survive a correction. A camera movement needs to communicate something more specific than “cinematic.”

For September, we rescored the roster across four disciplines with those demands in mind. The interesting result is not that certain models have moved ahead. It is how differently they get there, and how quickly a supposed weakness becomes irrelevant when you choose the right tool for the actual job.

THE NUMBERS NEED TO MEAN SOMETHING

Every capability is scored from 0 to 100. A score of 95 or above represents the best in the world at that particular dimension in this assessment. 90 to 94 is frontier territory. 85 to 89 is strong professional production capability. 78 to 84 means solid, but visibly behind the leaders. 70 to 77 means usable with workarounds. Below 70, the capability is weak or largely unsupported.

We use the whole scale because the missing capabilities are often the ones that determine whether you can finish.

A celebrated model without a meaningful edit mode can score in the 60s for editing. A world generator that takes minutes to respond can receive a 60 for latency while producing exceptional environments. Those are useful distinctions, not insults. Squeezing every tool into an 85-to-95 comfort zone makes the chart friendlier and the decision harder.

The scoring process starts with a vendor-page audit of the roster, with unresolved verification gaps disclosed below. Three independent research passes then assess every dimension. One is anchored in Artificial Analysis and Arena.ai Elo data sampled September 4 to 7, 2026. Another approaches the models as a post-production supervisor would, prioritizing control, reliability and failure modes. The third explicitly challenges inflated claims, known regressions and scores that are suspiciously generous.

Each published cell takes the median of those three assessments, followed by a review of naming and ordering.

That is an editorial method, not a laboratory certification. Video and image have public arenas to draw from. The 3D and realtime sections rely more heavily on vendor documentation and hands-on behavior. A number should make the judgment easier to inspect, not disguise the fact that judgment exists.

VIDEO: THE SHOT HAS TO SURVIVE THE NEXT SHOT

The video frontier is now three models deep, with a fourth making a particularly strong argument wherever continuity and camera direction matter.

Gemini Omni Flash, Google DeepMind’s June 30 API release, leads the sampled Artificial Analysis text-to-video arena and scores 94 for both quality and adherence. Its strongest production argument is edit-in-place: a client revision does not automatically require throwing away the entire generation.

That is meaningful control. It also comes with meaningful limits. Native output is 720p, clips cap at 10 seconds, and documented character drift across scene changes holds continuity at 85. Detail lands at 82. Being excellent at taking a note does not erase those delivery constraints.

Wan3.0, generally available from Alibaba since August 24, matches Omni Flash on quality and takes the section’s only 95 for editing. It can generate 30 seconds in one pass, work from documents and decks, use up to 20 reference assets, and regenerate selected intervals within a finished cut.

The attraction is straightforward: preserve what already works and intervene where it does not. Its 1080p ceiling, multi-character drift and reported in-frame text errors still need to be part of the decision.

MiniMax H3, released July 31, is the strongest image-to-video choice in the assessment and the most balanced video row. It delivers native 2K at 24fps, stereo audio and clips up to 15 seconds, alongside voice and motion transfer. Its brand-accurate text rendering is the strongest we have seen in video. Camera control remains at 87, because steering depends on reference video rather than explicit parameters.

Seedance 2.5, also released July 31, wins continuity at 94 and camera at 92. Its 3D white-model camera blockouts and support for up to 50 multimodal references give it a particular advantage when the scene has already been directed in your head and you need the system to follow.

The live API currently serves 480p and 720p, however. Quality lands at 91, detail at 83. Control and resolution are separate questions, and the matrix refuses to pretend otherwise.

The familiar defaults have become more specialized. Veo 3.1 remains the safest choice in this assessment for brand-safe 4K dialogue work inside Google Flow, despite placing 14th in the sampled Artificial Analysis arena. Runway Gen-4.5 retains distinctive value for colorist-led post through ProRes, HDR and 12-bit delivery. Its lack of native audio and reliance on a separate model for video-to-video leave it at 83 across the board.

Neither becomes useless because something newer scores higher. Their value depends on what happens after generation.

Kling 3.0 Omni scores 90 for camera control, with explicit pan, tilt, dolly, crane and rack-focus controls, plus native 4K at 60fps. Its arena position is less impressive than its marketing suggests. Luma’s Ray3.2 remains the keyframe and EXR specialist, with the section’s second-highest editing score at 89.

Sora 2 is absent. Its consumer app closed April 26, and its API is scheduled to shut down September 24. Availability is a production requirement, too.

IMAGE: BEAUTIFUL IS THE FIRST ROUND

An image generator can impress you once. A campaign tool has to keep working after the art director asks for changes.

That is where GPT Image 2.5, released September 8, establishes the clearest lead in the entire matrix. It scores 96 for quality, adherence and editing, with 93 for text. No other image model comes within five points of its editing score.

The Sunburst precision variant is the choice for multi-round campaign retouching without drift. Flare is the faster default. The difference is practical: an approved image should not become a negotiation with an entirely new image every time someone changes their mind.

The next tier is competitive for different reasons. Microsoft’s MAI-Image-2.6 scores 90 for quality, 91 for adherence and 91 for text, positioning it as a photoreal product and portrait workhorse. Grok Imagine Image 2.0 matches its quality score at roughly four cents per image and supports five-reference composites. Reve 2.1 offers native 4K, 91 for text, and per-element re-renders on approved layouts.

Google’s Gemini 3 Pro Image, also known as Nano Banana Pro, and Gemini 3.1 Flash Image, or Nano Banana 2, both score 89 for quality. Pro is the stronger choice for search-grounded infographics and multi-character identity locking, with 90 for consistency. Flash is the high-volume social variation engine, with 14-reference control.

Then there are the specialists, which is where blanket declarations about the “best image model” start becoming unhelpful.

Ideogram 4.0 leads text at 94. Recraft V4.1 leads style at 94, with native SVG output. Midjourney V8.2 Edit Model, released August 27, leads creativity at 95. Its text rendering sits at 72, and access remains web and Discord only, without an API.

For an artist developing an unusual visual language, that creativity score may matter more than automation. For a team generating a large set of tightly controlled commercial assets, the calculation changes. Neither brief needs permission from the other.

The lower end of the table shows the cost of standing still. FLUX.1 Kontext has fallen into the 70s, while Black Forest Labs now directs users toward FLUX.2 [klein]. Stable Image Ultra, at 74 for quality, occupies a legacy budget position. Adobe Firefly Image Model 5 scores 82 for quality and earns its place through commercial safety and IP indemnity rather than leading output.

The image you prefer and the tool your production needs may not be the same thing.

3D: A BEAUTIFUL MESH CAN STILL BE UNFINISHED WORK

The 3D section makes the distinction between appearance and usability especially difficult to ignore.

Meshy 7.1, released September 10, takes the section’s only 95 for mesh quality, with ultra geometry at 4096³ resolution and an integrated toolkit covering rigging, retexturing and auto-splitting. Hi3D 3.0 scores 92 for mesh, making it a strong collectible-grade close-up specialist. Rodin Gen-2.5 leads textures at 93, with 12K output and photoreal quad-topology hero assets.

Those scores describe different kinds of strength. They do not eliminate the next question: what does the asset need to do?

Dense-triangle generators including TRELLIS.2, Seed3D 2.0 and Hi3D 3.0 score between 84 and 92 for mesh, but only 72 to 76 for topology. A beautiful surface is not automatically a suitable structure for rigging or realtime use.

Read the topology and realtime-readiness columns instead, and the hierarchy changes. Meshy T2 scores 92 in both. Sloyd API scores 86 for topology and 90 for realtime readiness. Both have mesh scores in the 70s.

For the right production, that is a better trade. The asset that looks less spectacular in a comparison image may be the one that requires less intervention to do its job.

Seed3D 2.0 has a distinct simulation use case through URDF export into Isaac Sim with physics sets. TRELLIS.2 remains the open-source, self-hosted option with full PBR materials, though its editing score is 66.

Choose the asset’s destination before choosing its generator. Otherwise, you risk selecting the cleanup job with the prettiest preview.

REALTIME: A CONVERSATION IS NOT A WORLD

The realtime category needs a distinction that its name tends to conceal. A live voice agent and a generated environment are different products. Their strengths should not be flattened into a single idea of responsiveness.

gpt-live-1, introduced through OpenAI’s Live API on September 10, becomes the voice reference in this assessment, scoring 93 for quality and latency. It is full-duplex, meaning it can listen while speaking, with native turn detection and client-side delegation to your own tool pipeline.

It does not accept image or video input. Spatial context therefore sits at 70. A strong conversation does not imply visual understanding.

Gemini 3.8 Live, stable September 15, is the multilingual and multimodal choice, with 97-language support, live camera input and 88 for spatial context. Its Extended Thinking variant trades latency, scored at 84, for reasoning and the section’s highest interactivity score, 90.

gpt-realtime-2.1 remains the most capable tool-oriented voice agent in the table at 92 for tooling, supported by the most mature function-calling and telephony stack in the assessment. It remains the sensible pick for complex phone workflows involving multiple tools.

World generation asks for a different reading.

World Labs’ World API leads consistency at 95 and spatial coherence at 96, with splat and GLB exports suited to location scouting and previs. Generation is asynchronous and takes minutes, so latency scores 60.

That is not a contradiction. A system can produce an exceptionally coherent environment without producing it quickly enough for a live interaction. Whether the wait is acceptable depends on the work.

Genie 3 is useful for rapid mood-world prototyping, but the absence of a public API leaves tooling at 52. Legacy gpt-realtime now occupies a budget position in the 70s and primarily belongs in deployments that have not yet migrated.

The category name is not a substitute for reading the columns.

READ THE FINE PRINT BEFORE YOU QUOTE THE BIG NUMBER

Video and image scores blend the September 4 to 7 Artificial Analysis and Arena.ai sample with vendor documentation. They are not laboratory benchmarks. The 3D and realtime scores have no equivalent public arenas behind them and should be read as editorial judgments grounded in documentation and assessed behavior.

Four entries could not be fully verified against vendor pages at the time of scoring: Hunyuan3D 3.1, Seed3D 2.0, Tripo H3.1 and Midjourney Video V1. That limitation belongs beside the conclusions, not somewhere nobody will read.

Meta’s Muse Image and Muse Video, FLUX 3 Video and MAGI-2 qualified through arena results but were omitted to keep the table readable. They remain candidates for the next edition. Absence here is not a judgment of irrelevance.

THE BRIEF GETS THE FINAL VOTE

The matrix is a shortlist tool. Start with the section, identify the two or three capabilities your project cannot compromise on, and read those columns first.

A 94 for quality cannot turn a 10-second limit into a 30-second shot. A creativity score does not fix an unreadable product label. Exceptional mesh detail does not guarantee an asset is ready to rig.

This is why model loyalty is such a strange creative position. The point is to make the work you intended, with as much control over the result as possible. Sometimes that means choosing the model everyone is talking about. Sometimes it means choosing the less fashionable one because it gives your editor, animator or art director exactly what they need.

The artist’s judgment still matters. So does the unglamorous business of getting the file into the next person’s hands without an apology attached.

The model gets a score. You still have to make the film.

The STAGES × NAKID Generative Model Comparison Matrix is updated as the roster changes. This edition reflects the supplied September 2026 assessment.



Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.