IA 360
Current Affairs

Google launches Gemini 2.5 Flash Image, the “Nano Banana” model

Google reveals that the anonymous “Nano Banana” model was Gemini 2.5 Flash Image. It can edit photos through prompts, maintain characters across scenes and combine images for about $0.039 each through the API.

5 min read AI-generated Leer en español
Google launches Gemini 2.5 Flash Image, the “Nano Banana” model

On August 26, 2025, Google introduced Gemini 2.5 Flash Image for image generation and editing. The original source supports the documentary core of the event; selected examples and an API price do not measure fidelity, consistency or the total cost of a real workflow.

The launch identifies one of the anonymous models that had attracted the most attention on LMArena. Its main strengths are natural-language editing, combining multiple images and greater consistency when depicting the same character across different scenes.

Google’s anonymous model stood out on LMArena

“Nano Banana” had been circulating on LMArena for days, on a platform where users compare model outputs without initially knowing which system produced each image. The model reached the top spot in its image-editing leaderboard, according to Google.

These blind tests measure which results people prefer, but they are not a definitive evaluation. A model’s position depends on the participants, the prompts used and the types of images being compared. Even so, the strong performance suggested that a competitive system was behind the alias, particularly for modifying existing photos.

Google now presents it as an evolution of the image generation built into Gemini 2.0 Flash at the beginning of 2025. The goal is not just to create a scene from scratch, but to give users more control over an image through written instructions.

Edit a photo as if you were giving instructions to a person

Gemini 2.5 Flash Image lets users request specific changes in everyday language: replace the background, change the clothing, remove an element or transform the visual style without manually selecting each area. It also supports successive edits within a conversation, allowing users to refine the result step by step.

Another notable feature is blending multiple images. The model can take elements from different photos and bring them together in a single composition—for example, placing an object from one image into the setting of another or combining a person with a product.

Google also promises greater character consistency. This has been a common weakness of visual generators: when users requested new poses, outfits or settings, a person’s features could change so much that they became unrecognizable. The new model aims to preserve those features, making it easier to create image series, advertising materials, illustrated stories or variations of the same photo.

That does not amount to a guarantee of perfect identity preservation, however. Accumulated edits, small faces or complex compositions can introduce differences. Commercial work will still require reviewing every result and checking that the model has not altered important details.

It costs $0.039 per image through the API

Gemini 2.5 Flash Image is available to developers through the Gemini API and Google AI Studio, and to businesses through Vertex AI. The announced price is $30 per 1 million output tokens. Google estimates 1,290 tokens per image, putting the generation cost at $0.039 per image, or just under four cents. Document supporting the figure.

Producing 1,000 images would therefore cost about $39 in output tokens alone. The calculation does not include any additional cost from text or image inputs, which follow Gemini 2.5 Flash pricing. Document supporting the figure.

The combination of low cost and conversational editing points to high-volume uses such as e-commerce catalogs, ad variants, design prototypes and social media content. Consistency across images may be more valuable for these tasks than producing a single particularly eye-catching illustration.

SynthID helps trace AI-generated content

Google is incorporating SynthID, its invisible digital watermarking system, into images created or edited with the model. The signal is embedded in the file to help identify AI-generated content without necessarily placing a visible label on the image. In the Gemini app, Google also applies a visible mark to generated content.

This measure is especially important for a system capable of preserving faces while changing contexts, clothing or actions. The same flexibility that can be used to create campaigns or stylized photos can also be used to fabricate deceptive images.

SynthID provides a way to verify content within Google’s ecosystem, but it does not by itself resolve the question of digital content provenance. Platforms will need to integrate tools capable of detecting the watermark, and users will still have to assess an image’s source and context.

With Gemini 2.5 Flash Image, Google is competing less to generate the most spectacular illustration and more to turn visual editing into a conversational task. The decisive test will be whether it preserves identity and details reliably enough when businesses and creators incorporate it into repetitive workflows.

Turning the headline into a check

The first step is to freeze the system's identity. Google introduced Gemini 2.5 Flash Image for image generation and editing. A commercial name may cover different revisions, automatic routes and tools. A test record should preserve date, access mode, configuration, permissions and the full output. Without that snapshot, an improvement or failure observed today cannot rigorously be attributed to the version another person will use tomorrow.

Next, turn how to test editing, identity, provenance and cost before integrating the model into cases with acceptance criteria. Build a local sample containing easy, ambiguous, long and deliberately impossible tasks. Record the input, the information the model may consult and what outcome would count as sufficient. A vendor-selected demonstration shows possibility; a test set preserved by the user measures reliability.

Autonomy needs a permission ladder. Reading and proposing are not the same as editing, sending or buying. A safer setup begins with read-only access, requires a preview and reserves execution for explicit approval. It also keeps a log and a rollback path. Judge the model by the errors the surrounding system contains, not by the confidence of its plan.

What the record must preserve

Cost and quality must be measured together. A cheap answer that must be reviewed from scratch can cost more than a slower but verifiable one. Measurement includes waiting, retries, consumption, human oversight and the consequences of failure. selected examples and an API price do not measure fidelity, consistency or the total cost of a real workflow. That boundary turns the announcement into a testable hypothesis rather than a promise to be believed.

An evidence sheet separates four columns: what the source claims, what it shows, what it did not measure and what would change the conclusion. That discipline prevents an absence from becoming a promise and a condition from vanishing in summary. It also lets the story be updated without rewriting history from a later outcome.

Include a negative case before deciding. Find a situation where the system, rule, transaction or study does not meet the need and record the signal that would require stopping. Selected successes show that something can happen; the negative case reveals the boundary and lowers the cost of discovering it after deployment.

The skill that outlasts the announcement

A valid comparison preserves denominator and axis. It does not pit a point figure against an average, future capacity against installed capacity or a forecast against an observation. When two sources use similar language, reconstruct what they counted and over what period. If those differ, publish them as different measures instead of inventing a ranking.

The record should survive a version change. Keep URL, consultation date, document, configuration and decision. When new evidence appears, add it with its date and explain what it changes. That traceability prevents opposite errors: keeping an expired conclusion or pretending later information was known on the event date.

The transferable skill in this story is how to test editing, identity, provenance and cost before integrating the model. The procedure is short: name the document, preserve the date, fix the axis, find the condition and design a check that can fail. With those steps, a reader need not accept or reject the announcement by intuition; the decision follows a visible chain of evidence.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close