Google Launches Gemini 3.5 Live Translate: Real-Time Voice Translation Across 70+ Languages

VEKA

VEKA

June 11, 2026 (Updated: June 12, 2026) · 2 min read

Google Launches Gemini 3.5 Live Translate: Real-Time Voice Translation Across 70+ Languages

Beyond Text: Speaking Directly Across Languages

Machine translation spent a decade optimizing typed paragraphs. But spoken conversation demands something far harder: latency low enough to feel like dialogue, and prosody faithful enough that sarcasm doesn't arrive as monotone. On June 9, 2026, Google announced Gemini 3.5 Live Translate, a voice-to-voice mode for Gemini 3.5 that translates in near real time while preserving a speaker's pacing, pitch, and emotional color.

How It Works

Gemini 3.5 Live Translate is a streaming speech-to-speech audio model. It takes spoken audio in one language and emits translated spoken audio in another — across 70+ languages with over 2,000 language combinations. Two design decisions make it stand out:

  • Language auto-detection: The model detects 70+ languages without requiring manual source/target configuration. For call apps where you can't predict who speaks what, this removes an entire config step.
  • Prosody preservation: The output isn't flat TTS — it preserves the speaker's intonation, pacing, and pitch, making translations feel natural rather than robotic.

Unlike turn-by-turn systems that wait for a speaker to finish, Gemini 3.5 Live Translate generates speech continuously. It balances the trade-off between waiting for more context (better quality) and translating immediately (keeping sync with the speaker). The result stays just a few seconds behind the speaker throughout a session.

Where It Ships

The model is available through multiple channels:

  • Gemini Live API & Google AI Studio — for developers building custom integrations
  • Google Translate app — on both Android and iOS, for everyday use
  • Google Meet — language support jumps from 5 to over 70 languages with 2,000+ combinations

Ride-hailing service Grab is already testing the model for driver-passenger communication across Southeast Asia, where multilingual interactions are a daily reality.

Safety & Watermarking

All generated audio is tagged with an inaudible SynthID watermark — Google's deep fingerprinting technology that identifies AI-generated content without affecting audio quality. This ensures synthesized speech can be traced back to its origin, an important safeguard as real-time translation blurs the line between human and AI-generated voice.

Why It Matters

Language barriers have long been one of technology's most stubborn problems. Translation apps improved over the years, but conversations still felt awkward — you speak, wait for the translation, then wait again. Gemini 3.5 Live Translate aims to make multilingual conversations feel much closer to natural dialogue. For businesses with global teams, travelers, and developers building communication tools, this is a meaningful step toward a world where language is no longer a barrier to connection.

VEKA

VEKA

Author at VEKA

Share this article

// THREAD

Discussion