Google’s Gemini 3.5 Transcribe: What It Does and Who It’s For

Gemini 3.5 Transcribe is Google’s newly released public-preview speech-to-text model, designed specifically for both low-latency live captioning and prerecorded audio processing. The model includes language auto-detection, custom vocabulary support, smart formatting, speaker diarization, and word-level timestamps, positioning it as a purpose-built alternative to general-purpose AI models being prompted to handle audio tasks.

Currently in public preview, the model comes in two distinct versions: a file-based model for prerecorded audio and a separate live model optimized for real-time streaming. Google estimates pricing at roughly $0.005 per audio minute for file processing and about $0.009 per minute for live transcription.

What Makes Gemini 3.5 Transcribe Different

Unlike a general AI model simply asked to analyze audio, Gemini 3.5 Transcribe is a dedicated speech-to-text system with separate endpoints built specifically for live WebSocket streaming versus prerecorded file processing. This separation lets developers choose between prioritizing low latency or richer transcript metadata, depending on their specific use case.

The model automatically detects more than 85 languages and can follow code-switching within a single session, meaning it can handle speakers who shift between languages mid-conversation. Custom vocabulary biasing further helps improve accuracy around names, acronyms, product terminology, and specialized industry jargon.

One of its most distinctive capabilities is Smart transcription. Rather than preserving every spoken sound exactly as uttered, this feature can automatically remove filler words, repetitions, and speech disfluencies, resolve self-corrections mid-sentence, add proper punctuation, and normalize spoken values like currency amounts or identification numbers into clean, formatted text.

That kind of cleanup works well for everyday dictation, but it isn’t appropriate for every use case. Legal, journalistic, research, support, compliance, and evidentiary workflows often require verbatim transcripts that preserve hesitations, false starts, and exact spoken wording, meaning teams working in these areas should carefully evaluate whether Smart transcription’s cleanup features are actually appropriate for their specific needs.

Who Should Use Gemini 3.5 Transcribe

The right fit ultimately depends on the specific task at hand rather than simply the length of the feature list. For voice interfaces and live captioning, the model can stream microphone or call audio directly into interactive applications with sub-second transcription latency.

For meetings and call analytics, the file-based model processes recorded audio complete with speaker labels, timestamps, custom terminology support, and clean formatting. For polished dictation, Gemini 3.5 Transcribe can turn natural, filler-heavy speech and mid-sentence corrections into clean, readable text suitable for drafts, messages, or voice commands.

The model also supports multilingual products through its language auto-detection and accent handling, while domain-specific transcription needs benefit from custom vocabulary biasing toward product names, acronyms, people, locations, and specialized terminology.

Core Features Worth Knowing

Beyond its live streaming and prerecorded processing capabilities, the model supports speaker diarization, capable of attributing audio segments to as many as eight distinct speakers, though Google notes that attribution accuracy for three or more speakers remains experimental at this stage.

Word-level timestamps are also available for prerecorded transcripts, marking precise start and end offsets for individual words, useful for search functionality, playback syncing, video clips, and subtitle generation. The model additionally handles formatting and normalization automatically, adding proper capitalization and punctuation while converting spoken numbers or monetary values into clean, formatted text.

Notably, the same underlying model already powers voice experiences across several Google products, including the Gemini app on macOS and other announced Google surfaces.

Pricing and Access

Gemini 3.5 Transcribe offers a free developer tier alongside token-based paid pricing. Google estimates a blended cost of approximately $0.005 per audio minute for prerecorded transcription and $0.009 per minute for live transcription, though actual costs depend on audio-input and text-output tokenization rather than these flat estimates alone.

The free developer tier provides evaluation access through Google AI Studio, though it’s worth noting that free-tier content can be used to improve Google’s products and may be reviewed by human evaluators, making it unsuitable for sensitive or confidential recordings. Paid API usage, by contrast, is not used for product improvement under current developer terms, though Google still retains prompts and responses for a limited period for abuse prevention and legal compliance purposes.

Strengths and Limitations to Consider

The model’s purpose-built design for both live and file-based use cases avoids forcing a single endpoint to handle fundamentally incompatible latency and metadata requirements. Its broad language support, competitive per-minute pricing, and combination of diarization with word timestamps in one low-cost API make it a compelling option for large-scale transcription workloads.

That said, several limitations are worth understanding beforehand. Since the model remains in public preview, its behavior, limits, and pricing could still change before general availability. Live sessions are currently capped at 10 minutes and don’t support speaker diarization or word-level timestamps, while prerecorded requests support up to one hour, though enabling diarization or timestamps reduces that limit to 30 minutes.

Custom vocabulary accepts up to 1,000 terms, though Google recommends keeping lists closer to 100 focused terms for optimal accuracy. Transcripts can also still misrecognize names, numbers, accents, overlapping speech, and quiet or noisy audio, and teams should remain aware that polished formatting can sometimes make errors appear more authoritative than they actually are.

Given these considerations, appropriate consent, data retention policies, and jurisdiction-specific privacy controls remain essential whenever recording, transcribing, or storing audio involving real people.


Source: This article is based on information from Google’s official Gemini documentation and The Rundown AI.

AI News

Leave a Reply

Your email address will not be published. Required fields are marked *