Google pairs fast images and conversational video with Nano Banana 2 Lite and Omni Flash
Nano Banana 2 Lite arrives as a stable, lower-cost image model; Gemini Omni Flash offers conversational video but remains in preview.
On June 30, 2026, Google made Nano Banana 2 Lite and Gemini Omni Flash available to developers, two tools that fit into the same creative workflow even though they are not at the same stage of maturity. Nano Banana 2 Lite is built to generate and edit images quickly at a lower cost; Gemini Omni Flash turns text, images, and clips into video and lets users revise it with natural-language instructions.
The combination is easy to picture: generate many image variations, choose one, and use it as a reference for a clip. But the meaningful part of the announcement is not a flashy demo. It is the way it separates two needs that are often mixed together in generative media: high-volume visual iteration and video editing with continuity.
An image option for rapid iteration
Nano Banana 2 Lite is the product name for Gemini 3.1 Flash Lite Image, available as a stable model in the Gemini API. Google describes it as the fastest and least expensive member of its image family. Its announcement gives a text-to-image time of four seconds and a price of $0.034 per 1K-resolution image.
That makes it suitable for visual drafts, interactive interfaces, and applications that need to create many images without letting latency or per-request budgets grow too much. Google recommends the version as a replacement for Gemini 2.5 Flash Image, the first Nano Banana release.
That efficiency comes with trade-offs. The documentation says Lite is not optimized for working with multiple reference images or for sequential editing across many turns. It supports only 1K output and does not include some of the more ambitious functions of higher-tier models, such as search grounding. For a complex composition, a campaign requiring strong brand consistency, or a final 4K image, the same family offers other options.
Conversational video, still in preview
Gemini Omni Flash, listed in the API as gemini-omni-flash-preview, supplies the other half of the workflow. Google describes it as a model for generating and editing video from a combination of text, image, and video inputs, then refining the result through conversation. Its model catalog keeps it in preview rather than stable status.
The announced price is $0.10 for each second of generated video. At this stage, generations are ten seconds long. That does not rule out uses such as product clips, demonstrations, social posts, or scene prototypes, but it does mean that a longer production has to be planned as a sequence of shots rather than one request.
Google also lists specific limits: the API does not yet accept audio references or scene extension; reference videos up to three seconds appear in the schema but are not yet processed correctly; and character consistency can break when scenes change or the camera pans. Those details are more useful for choosing an integration than a broad promise of “AI video.”
Chaining tools without giving up control
The combined workflow is straightforward: Nano Banana 2 Lite produces a fast visual starting point, and Omni Flash animates or transforms it. Google offers examples in interior design, ecommerce, and scenes based on a photo. Its Interactions API can preserve session history for several consecutive edits, but a team moving toward production should test consistency, timing, and costs on its own material.
Images made with Nano Banana and content from these models include a SynthID watermark, according to Google. That is a useful provenance signal, but it does not replace checking rights in the input material, reviewing the output, or labeling AI use appropriately when the context requires it.
The practical choice is clear: Nano Banana 2 Lite fits cases that need many fast, inexpensive images; Omni Flash is for exploring conversational video generation and editing. The latter should still be treated as a preview. They are not two magic buttons, but components with different limits for building a faster multimedia workflow.
Sources for this piece
This piece draws on 4 primary source(s), gathered during reporting.
This article was produced with artificial intelligence under human editorial oversight.