Gemini 3.5 Live Translate: near real-time voice translation
Google DeepMind launches an audio model that translates speech to speech in more than 70 languages while preserving the speaker's tone and rhythm. It's rolling out to Google Translate, Meet, and developers via the Live API.
Google DeepMind has unveiled Gemini 3.5 Live Translate, an audio model that translates speech from one language to another continuously, without waiting for the speaker to finish a sentence. It automatically recognizes more than 70 languages and reproduces the translated voice while preserving the intonation, rhythm, and tone of the original speaker. The company began rolling it out on June 9, 2026, in Google Translate, Google Meet, and through the API for developers.
The promise is concrete: a conversation between two people who don't share a language, with a delay of just a few seconds, instead of the awkward pauses typical of automatic translators. Google frames this within the two-decade history of Google Translate, a service that, according to the company, already translates more than a trillion words a month for billions of users.
What's different from what already existed
The technical difference lies in how it processes audio. Standard voice translators work turn by turn: the system listens to a complete sentence, translates it, and only then plays it back. The result is accurate but mechanical, with long silences that break the flow of any real conversation.
Gemini 3.5 Live Translate works differently. It generates translated speech continuously, in a stream, while the other person keeps talking. The model has to strike a delicate balance: waiting a bit longer provides more context and improves translation quality, but translating immediately keeps pace with the speaker. According to DeepMind, the system stays only a few seconds behind the speaker throughout the entire session.
The other leap forward is that it preserves the voice's characteristics: intonation, speed, and tone. In practice, this means a question sounds like a question rather than a flat statement — something classic translators tend to lose along the way. The model also handles multilingual input without manual configuration and is designed to work in noisy, unpredictable environments.
Where it's rolling out, and when
Google is deploying the tool on three fronts with different levels of access:
- Developers: in public preview through the Gemini Live API and Google AI Studio.
- Businesses: in private preview this month within Google Meet, for select Google Workspace customers.
- General users: in the Google Translate app for Android and iOS, with a global rollout.
The Google Meet case offers the clearest numbers for how big this leap is. Voice translation on the platform went from supporting just five languages to more than 70, and from translating only to and from English to enabling more than 2,000 language combinations within a single meeting. It's a shift in scale that turns a niche feature into something genuinely usable for truly multinational teams. The wider rollout will arrive later this year.
In the Google Translate app, all it takes is plugging in a pair of headphones to use live translation. On Android, Google is also introducing a "listen mode": you hold the phone up to your ear as you would for a regular call, and the translation comes straight through the phone's own earpiece, with no external headphones needed. The company positions this for everyday situations — a guided tour in another language, for instance — where you want to hear the translation discreetly and don't have headphones on hand.
A partner ecosystem for real-world applications
The API route is likely the one with the most industrial potential. Voice infrastructure platforms such as Agora, Fishjam, LiveKit, Pipecat, and Vision Agents are already integrating the Gemini Live API. These companies handle the complex part — real-time audio streaming — so other developers can build their own voice translation apps without having to solve that technical plumbing themselves.
Among the use cases Google highlights is Grab, the Southeast Asian super-app, which is testing the model to help drivers and passengers who speak different languages understand each other at pickup points. According to the company, its users make more than 10 million voice calls a month through the platform. Other companies, including CJ ENM and LiveKit, have shared positive feedback focused on quality, accuracy, and low latency.
The intended uses go beyond one-on-one conversation: live interpretation for multilingual calls and meetings, classroom instruction, simultaneous dubbing, and broadcasts.
The invisible watermark
All audio generated by the model carries SynthID, Google DeepMind's watermark. It's an imperceptible signal embedded directly in the audio output that allows content to be identified as AI-generated. In the realm of synthetic voice — where cloning and impersonation are an obvious risk — that traceability is a meaningful safeguard, though it's worth remembering that its effectiveness depends on accessible verification tools existing and on other actors not being able to easily strip the watermark out.
What to watch
Google hasn't published comparative accuracy metrics or exact latency figures beyond "a few seconds behind," so much of the real-world assessment will depend on how it performs in users' hands and in less common languages, where translation models tend to struggle. Preserving tone is great, but a translation error delivered in a natural, convincing voice can be more misleading than one in a robotic voice — precisely because it sounds human.
The launch comes at a time when other major tech companies are also investing in real-time voice translation, though Google doesn't explicitly name specific competitors in the announcement. Google's edge is distribution: Translate, Meet, and an API open to third parties give it three distinct entry points, from the tourist holding a phone to their ear to the corporate meeting with participants speaking a dozen languages. The bar is no longer just translating well, but translating fast without breaking the flow of conversation. That's the territory Gemini 3.5 Live Translate aims to claim.
This article was produced with artificial intelligence under human editorial oversight.