OpenArt Launches Creative AI Arena That Splits Model Rankings by Job
The new public benchmark separates filmmaking, design, editing and lip sync. Its first results support that approach—but its undisclosed active prompts leave users with a trust question.
Listen to this story
The audio brief
Story brief
3 key pointsOpenArt has launched Arena, a public benchmark that ranks image and video models by creative workflow rather than one universal score. Seedance 2.5 led overall video at 1,081, while Wan 3.0 narrowly won video editing at 1,034; Seedream 5.0 Pro led overall image at 1,010, with GPT Image 2 stronger in graphic design and editing. Arena uses blinded pairwise judging and Bradley-Terry aggregation, but undisclosed voter...
- 01
Seedance 2.5 topped overall video and four of five specialized video boards supplied at launch.
- 02
Wan 3.0 beat Seedance 2.5 in video editing by one point, within a difference OpenArt says may not be meaningful.
- 03
Seedream 5.0 Pro led overall image, while GPT Image 2 ranked first in graphic design and image editing.
OpenArt’s new Arena promises a more useful answer than a single “best model” score: choose an AI image or video model for the creative job at hand. Its launch rankings do show different winners across tasks, but the benchmark will keep some active prompts private—putting transparency at the center of whether creative teams treat it as a decision tool.
OpenArt Arena is a public benchmark for AI image and video generation. Rather than collapse performance into one ranking, it publishes separate leaderboards for filmmaking, advertising, graphic design, e-commerce, motion design, video editing, lip sync and overall performance.
A benchmark built around the production task
The premise is straightforward: a model that handles cinematic lighting and camera movement well may not be the best fit for an ad that needs readable text, accurate logos and faithful product placement. OpenArt generates outputs from curated prompts for each board, then presents unlabeled pairs to evaluators for side-by-side choices.
Those preferences are aggregated with the Bradley-Terry model, a statistical method that estimates an entrant’s chance of winning from head-to-head comparisons. OpenArt planned to draw on a Creative Expert Council and roughly 800 to 1,000 additional “tastemakers,” though it did not disclose how many people completed the launch voting.
Seedance 2.5 ranked first on the overall video board, ahead of Wan 3.0 at 1,004 and Seedance 2.0 at 1,000.
Wan 3.0 placed first for video editing by one point over Seedance 2.5, a gap OpenArt’s confidence intervals caution against treating as a meaningful quality difference by itself.
Seedream 5.0 Pro led the overall image ranking, 10 points ahead of GPT Image 2.
The launch results make the case—and limit it
The video table produces a strong general leader. Seedance 2.5 topped the overall board and ranked first on four of the five specialized video leaderboards supplied for launch. But Wan 3.0 edged it in video editing, illustrating why an overall winner may not settle a team’s choice for a specific workflow.
The image boards are more divided. GPT Image 2 ranked first for graphic design and image editing, while Seedream 5.0 Pro led the film-oriented imagery and e-commerce boards, as well as the overall image ranking. That pattern is the practical value of task-specific comparison: it can direct users toward candidates rather than declare one universal champion.
The unanswered question is what the judges saw
OpenArt says it will publish its methodology and part of its prompt sets, while withholding some active prompts to make it harder for model makers to tune directly for the test. That protects against a static benchmark becoming a target. It also means users cannot fully inspect whether the hidden prompts represent the kinds of creative work they need to do.
For now, Arena is most useful as a structured starting point, not a final procurement verdict. The close video-editing result also shows why teams should read the task board and its confidence intervals, rather than treat small score differences as proof that one model is materially better.
Sources
- venturebeat.comWhat's the best AI model for graphic design, video ads, lip sync and more? OpenArt's new Arena offers leaderboards for different media jobs
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.