SyncAI.news, a Varaisys broadcasting
Gemini 3.5 Transcribe vs OpenAI’s GPT-Transcribe
SO

Shittu Olumide

· 1 min read

EngineeringKDnuggets

Gemini 3.5 Transcribe vs OpenAI’s GPT-Transcribe

Google shipped Gemini 3.5 Transcribe on August 26, 2026, and the timing makes it a genuinely useful comparison. OpenAI had released its own current flagship transcription model, GPT-Transcribe, just four weeks earlier, on July 28, 2026. Two labs, two new transcription models, released close enough together that comparing them actually means something right now instead of stacking one model generation against another.

Both companies split their offering the same way too — one model built for real-time streaming, one built for pre-recorded audio — which makes the comparison unusually apples-to-apples. Here's how each got to where it is, a real use case and working code for both, and a side-by-side on the numbers that actually matter.

Gemini 3.5 Transcribe replaces Chirp 3, Google's previous transcription model, and the improvement Google is leaning on hardest is speed: a 70% improvement in time-to-final-transcription over Chirp 3, alongside better accuracy. It ships as two distinct model IDs rather than one general-purpose endpoint: gemini-3.5-transcribe-live for continuous, sub-second-latency streaming through the Live API, and gemini-3.5-transcribe for pre-recorded audio, meetings, call logs, and similar, through the Interactions API.

The real numbers, as measured by Artificial Analysis and cited directly in Google's announcement: a 4.0% word error rate (WER) for streaming use and 2.6% for non-streaming. On the FLEURS multilingual benchmark specifically, Google reports 5.50% WER streaming and 5.04% non-streaming — worth noting as a separate, harder benchmark rather than mixing the two numbers together.

OpenAI's GPT-Transcribe

Let's take a quick look at some use cases.

Using Gemini 3.5 Transcribe for a Multi-Speaker Meeting

Consider a real scenario where the built-in diarization actually earns its keep: transcribing a recorded three-person meeting and getting back who said what, not just a wall of undifferentiated text.

Original source

This story was published by KDnuggets and written by Shittu Olumide. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on kdnuggets.com

Similar News