Artificial Analysis Rebuilds Image Editing Arena With Different Leaders by Task
The revised benchmark makes model selection a task-by-task decision: its overall winner does not lead every kind of edit, and listed API prices vary sharply.
Listen to this story
The audio brief
Story brief
3 key pointsThe rebuilt benchmark makes image-editing model selection more task-specific: MAI-Image-2.6-Preview leads overall, but GPT Image 2 (high) wins object edits and framing, while Seedream 5.0 Pro leads identity preservation. The 17 boards aggregate blind human preferences using Elo-style Bradley-Terry scoring, so they indicate perceived output quality rather than editing accuracy. For teams, the practical consequence is...
- 01
MAI-Image-2.6-Preview scores 1,284 Elo from 13,492 samples, only 25 points ahead of fifth-ranked MAI-Image-2.5.
- 02
GPT Image 2 (high) leads Object-Level Edit and Composition and Framing despite ranking fourth overall at 1,259 Elo.
- 03
Seedream 5.0 Pro tops Identity-Preserving Edit, covering likeness, outfits, accessories, poses, and expressions.
Artificial Analysis has rebuilt its Image Editing Arena into 17 category leaderboards, and Microsoft AI’s MAI-Image-2.6-Preview now leads the overall table. The change puts a practical limit on any simple “best model” verdict: the new results show different systems leading object edits, framing, identity preservation, and broader transformations.
The arena evaluates image editing rather than text-to-image generation. Models receive an existing image and an instruction, then return a modified version. Artificial Analysis has organized those tests across seven kinds of editing actions and 10 use cases, yielding 17 category boards. A planned Reference to Image benchmark is intended to separately assess generating new images from reference inputs.
The overall score comes from blind preferences
The ranking is built from blind user comparisons of two edited outputs generated from the same input image and instruction. Those preferences are converted into an Elo-style score, a rating format in which a higher number means an output was preferred more often. Artificial Analysis uses Bradley-Terry maximum-likelihood estimation for the underlying comparison model and refreshes prompts monthly.
That method makes the table a measure of aggregated human preference in this arena, not a universal measure of editing accuracy. MAI-Image-2.6-Preview holds 1,284 Elo from 13,492 samples; MAI-Image-2.5-Pro follows at 1,272, Reve 2.1 at 1,260, GPT Image 2 (high) at 1,259, and MAI-Image-2.5 at 1,257.
MAI-Image-2.6-Preview is ranked first overall with 1,284 Elo based on 13,492 samples.
GPT Image 2 (high) ranks fourth overall at 1,259 Elo.
Tencent’s HunyuanImage 3.0 Instruct leads the open-weights entries at 1,223 Elo.
Category boards change the buying decision
Microsoft’s new leader ranks first on seven of the 17 category boards. But GPT Image 2 (high) leads Object-Level Edit and Composition and Framing, categories covering precise local changes and changes to an image’s spatial framing. Seedream 5.0 Pro ranks first for Identity-Preserving Edit, which tests whether faces, likeness, outfits, and accessories survive changes in pose or expression.
Three distinct strengths in the rebuilt arena
- Overall leader and first on seven category boards: MAI-Image-2.6-Preview.
- GPT Image 2 (high): leader for object-level edits and composition or framing changes.
- Seedream 5.0 Pro: top-ranked for identity-preserving edits.
The result is a routing problem for teams that edit different kinds of assets. A system used for marketing transformations, UI work, character images, and tightly controlled object changes may benefit from sending each job to the model strongest in that category, rather than treating the overall leader as a universal default.
Quality is only one operating variable
The leaderboard shows a substantial price spread among prominent options: MAI-Image-2.5-Flash is listed at $20 per 1,000 images, compared with $211 for GPT Image 2 (high). These are not direct cost-of-work comparisons; Artificial Analysis defines them as creator API prices for 1,000 images at 1024-by-1024 resolution using each model’s default settings.
The table also identifies HunyuanImage 3.0 Instruct as its leading open-weights entry, ranked 13th overall with 1,223 Elo.
Sources
- artificialanalysis.aiImage Editing Leaderboard - Top AI Image Models | Artificial Analysis