Google has released Gemini 3.5 Transcribe in public preview, giving developers separate APIs for real-time voice interaction and recorded-audio processing. The split matters because a live assistant needs speed, while a call archive needs speaker labels and timestamps that can be searched or routed downstream.
One model, two operating modes
For interactive applications, the Live API exposes gemini-3.5-transcribe-live as a continuous, bidirectional stream with Google-stated sub-second latency. That is the route for voice agents and live captioning, where the system must keep accepting audio as it returns results.
The Interactions API uses a different model name, gemini-3.5-transcribe, for meetings, call logs, and other completed recordings. Its added outputs are word-level timestamps and attribution for as many as three speakers; Google labels support beyond three speakers experimental.
The API choice follows the job
01gemini-3.5-transcribe-liveLive API
Google positions the Live API for continuous bidirectional transcription in interactive voice applications, with stated sub-second latency.
02gemini-3.5-transcribeInteractions API
Google positions the Interactions API for completed audio and includes speaker attribution and word-level timestamps.
Recognition is only part of the output
The model is designed to do more than render utterances word for word. Google says it can remove filler words, resolve a spoken self-correction, format the result, and adapt to a supplied custom vocabulary. That makes the output a cleaned-up working draft, not necessarily a verbatim record.
Google also emphasizes terms that frequently break automated workflows: postal codes, order IDs, and other alphanumeric entities, including in noisy recordings. It says the model automatically detects and transcribes more than 85 languages, and can handle regional accents and dialects.
Google is distributing the capability across several surfaces
- Developers can access the public preview through the Gemini API in Google AI Studio and Google Antigravity, while enterprises can use the Gemini Enterprise Agent Platform preview.
- Consumer availability includes the Gemini app on macOS in English and Rambler on Android in selected countries and languages; Chrome support is planned.
- On Android, Rambler uses the transcription system for formatted dictation and voice-led edits such as correcting misspellings or changing writing style.
Performance claims meet an incomplete buying picture
Google says Artificial Analysis measured average word error rates of 4.0% for streaming and 2.6% for non-streaming use. Google also says time to final transcription improved 70% over Chirp 3, its prior transcription model. Those results distinguish the two modes on accuracy, but they remain performance figures presented in Google’s launch material.
The launch does establish model identifiers and API paths, addressing the practical question of how each mode is called. But Google’s pricing page does not list a price for Gemini 3.5 Transcribe. For a preview aimed at voice agents, captioning tools, and post-call analytics, that leaves the cost comparison unresolved.
Reader comments
Newest comments first. Replies stay oldest first.