Google Releases Nano Banana 2.1 With Lower Image Prices and Stronger Editing Claims
Google’s evaluations favor the new model over its predecessors, including Pro. A small hands-on comparison shows why those scores are not a universal quality verdict.
Google’s Nano Banana 2.1, built on Gemini 3.6 Flash and documented in an October 6, 2026 model card, makes image generation substantially cheaper: The Decoder lists 1K output at 3.36 cents and 4K at 7.56 cents, about half Nano Banana 2’s prices. Google also reports stronger preference and editing scores, especially with Thinking enabled, but those are company evaluations rather than independent rankings. For teams choosing image APIs, the price drop may make iteration cheaper; text rendering, edit fidelity and a hands-on comparison with Pro remain reasons to test outputs before switching.
01
With Thinking, Google reports a text-to-image preference Elo of 1,050 versus 990 for Nano Banana 2 and 935 for Nano Banana Pro.
02
Thinking raised Nano Banana 2.1’s overall preference score from 1,015 to 1,050 and its general-editing score from 980 to 1,026.
03
The model supports 1K, 2K and 4K output, and Google says it can use 14 reference images while maintaining up to four characters and ten objects.
Google has released Nano Banana 2.1, an AI model for creating and editing images, with stronger results in its own evaluations. The Decoder reports image-generation prices roughly half those of Nano Banana 2. The release brings cheaper output and improved editing claims across Google’s products, though neither company scores nor lower prices settle which model produces the best finished image.
Lower prices, multiple ways in
Google’s model card, published October 6, 2026, says Nano Banana 2.1 is built on Gemini 3.6 Flash. It accepts text and images and can return both images and text. Its listed distribution channels include the Gemini app, Google AI Studio, the Gemini API, Search AI Mode, Google Ads, Flow and Stitch.
The model supports 1K, 2K and 4K image resolutions. According to Google, it can work with up to 14 reference images while maintaining consistency for as many as four characters and ten objects. It also connects to Google Search and offers minimal, medium and high Thinking settings that affect image quality.
The Decoder’s pricing figures put the reduction at roughly half for both 1K and 4K images, rather than only the lowest-resolution option. Its table lists a 2K image at 5.04 cents. For comparison, Nano Banana Pro is listed at 13.40 cents for a 1K image.
The scores favor 2.1; one image favors Pro
Google evaluates image generation and editing separately. Its approach includes public benchmarks and internal tests, with human raters comparing outputs side by side. Those preferences produce Elo scores—a relative rating within the evaluation—while an automated evaluator assesses factuality. The tested tasks span infographic design, character editing, stylization and combining multiple reference images.
With Thinking enabled, Nano Banana 2.1 scored 1,050 on overall text-to-image preference, against 990 for Nano Banana 2 and 935 for Nano Banana Pro. Google also reports a multi-character consistency score of 1,106, compared with 978 and 1,011, respectively. These are Google’s evaluation results, not an independent ranking of finished-image quality.
Thinking also changes the comparison within the new model. Without it, Google reports an overall preference score of 1,015 and a general-editing score of 980. With it, those figures rise to 1,050 and 1,026. The card therefore presents two performance profiles for Nano Banana 2.1, not a single result that applies to every setting.
A hands-on example from The Decoder illustrates the gap between scores and a particular image. It used a deliberately awkward prompt involving a horse riding an astronaut. The outlet judged Pro’s result more realistic and natural in color and proportions; Nano Banana 2.1 made the horse look more like a pony. That narrow comparison is not a comprehensive benchmark.
Editing gains still leave visible failure modes
Google positions the model for precise image creation, repeated revisions, posters and intricate diagrams. Its own limitations qualify those uses. Small text can remain blurry, particularly at 1K, and long paragraphs are difficult to render. Character appearance is not always preserved between a reference image and the generated result.
Targeted edits can follow instructions only partly. Edits guided by a mask or doodle may also retain marks from the input.
A subject’s original pose can persist even when an edit calls for a structural change. Google describes these instances as rare.
The model can confuse left and right, and remains limited in advanced world knowledge, three-dimensional reasoning and factual accuracy.
Reader comments
Newest comments first. Replies stay oldest first.