Google just dropped Gemini 3.5 Transcribe, and it's their most accurate speech-to-text model yet, built to handle real conversations instead of just clean studio audio. - Word error rate down to 4.0% for streaming and 2.6% for pre-recorded audio, measured by Artificial Analysis - That's a big jump from their old model, Chirp 3. Time to final transcription is 70% faster - Handles up to 85 languages and picks up on accents and dialects automatically - Cleans up your speech as you talk. Cuts filler words like "um" and "ah" and fixes self-corrections like "let's meet Tuesday, no wait, Wednesday" - Can identify up to 3 speakers in recorded audio with timestamps, which is huge for meeting notes and call analytics - It can also call other Gemini models to do stuff like generate images or analyze files, all triggered by voice - Already live in things like Rambler on Android and the Gemini app on macOS, and rolling into Chrome soon - Developers can plug it into their own apps now through the Gemini API and Google AI Studio If you're building any kind of voice agent, meeting notes tool, or automation that starts with "listen to this audio," this is worth testing against whatever you're using now. I still use Deepgram Nova 3 or Flux for my voice agents. But I might start using Gemini 3.5 Transcribe if I like the results. Will ultimately come down to latency. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe