Sources
Learn the craft
Step-by-step guides on prompting, styles, and getting the most out of AI image generation.
Read the guides
Theo turns AI news into things you can actually try in tonight's session.
Step-by-step guides on prompting, styles, and getting the most out of AI image generation.
Read the guidesPick a companion and get their take on this story
Google has launched Gemini 3.5 Transcribe, a dedicated speech-to-text model that automatically removes filler words, detects specialized jargon, and supports more than 85 languages — changes that matter most if you're already narrating prompts, logging creative sessions by voice, or building audio-driven pipelines.
The headline capability isn't the language count — it's the automatic cleanup. If you dictate prompt ideas, record reference notes while reviewing renders, or use voice-to-text to feed a downstream text model, raw transcripts are usually cluttered enough that you spend real time fixing them before they're usable. Gemini 3.5 Transcribe handles that pass for you.
A concrete scenario: a concept artist who records a five-minute voice memo describing a scene's lighting, mood, and character details can now pipe that audio directly into a prompt-refinement step without manually scrubbing the transcript. The jargon detection means terms like "subsurface scattering" or "chromatic aberration" are less likely to get mangled into phonetic guesses.

Gemini 3.5 Transcribe is the latest addition to Google's expanding Gemini model family.
Image: The Verge / The Verge AI
According to The Verge, Gemini 3.5 Transcribe follows 3.5 Live Translate — a real-time translation model — while the more widely anticipated Gemini 3.5 Pro remains unreleased. That sequencing tells you something about Google's near-term priorities: audio understanding is shipping ahead of the flagship reasoning upgrade.
For creators, this means the Gemini audio stack is getting meaningfully better before the core generation model catches up. If you're evaluating which provider's API to route voice-driven workflows through, the transcription quality gap between Gemini and competitors just narrowed — or widened in Google's favor, depending on your current setup.
Automatic jargon detection is the quieter win here. Generic speech-to-text models trained on broad corpora often stumble on technical vocabulary — rendering terms, model names, or platform-specific language that doesn't appear in everyday speech data. Gemini 3.5 Transcribe is designed to recognize specialized language in context rather than forcing you to maintain a custom vocabulary list or correct the same mistranscription every session.
If you batch-process audio logs from a creative session — say, timestamped notes on which generations worked and why — cleaner jargon handling means those logs are actually searchable and useful downstream, not just a pile of phonetic approximations.
The model already powers Gboard's Rambler dictation feature and is expanding to Chrome and additional Google surfaces. That staged rollout suggests the API access for developers and third-party tools will follow, though Google hasn't given a specific timeline for broader availability. Creators building voice-to-prompt tools or audio-annotation pipelines should watch the Google AI developer channels for API access details.
For anyone already experimenting with voice-driven prompt workflows in the Charmloop generator, pairing a clean transcription layer with your existing process is a low-friction upgrade worth testing as access opens up. And if you're newer to building prompt pipelines, the guides section covers structuring prompts from rough notes — the kind of workflow a good transcription model slots directly into.