Google previews Gemini 3.5 Transcribe: fifth on the independent leaderboard it cites
Speech recognition is back at the centre of the contest between model providers, and it is the natural front door for voice assistants and agents. Against that backdrop, on 26 August 2026 Google announced on its official blog the public preview release of Gemini 3.5 Transcribe. The model, which aims to replace the Chirp line, does not simply convert audio word by word: it cleans it up and formats it on the spot. In the company's own words: "Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text". The technology is available in public preview through the Gemini API in Google AI Studio and on the Gemini Enterprise Agent Platform, shipping as `gemini-3.5-transcribe` for recorded audio (Interactions API) and `gemini-3.5-transcribe-live` for real-time interactions (Live API).
Google presents the model as its most accurate speech-to-text tool so far, but independent measurements paint a more layered picture. Google reports an error rate measured by Artificial Analysis of 2.6% in non-streaming mode and 4.0% in streaming. On that same organisation's public AA-WER leaderboard, consulted on 29 August 2026, the model sits in fifth place with a word error rate (WER) of 2.6%, tied with two other models and behind Fun-Realtime-ASR-preview (1.7%), ElevenLabs' Scribe v2 (2.2%), Microsoft Azure's MAI-Transcribe-1.5 (2.4%) and Smallest AI Pulse Pro (2.4%). Artificial Analysis notes in its methodology that a lower AA-WER means a more accurate transcript, calculated as an average weighted by audio duration over roughly 8 hours of test material. On the multilingual FLEURS benchmark, Google claims a WER of 5.04% in non-streaming and 5.50% in streaming, but those are measurements taken over a selection of languages and local variants that is not fully specified, which rules out a direct comparison. The company also claims a 70% improvement in processing time over Chirp 3, yet there is no independent verification and Google does not publish the details of the comparison.
The stated features include support for more than 85 languages with automatic detection, removal of fillers such as "um" or "uh", and speaker attribution for up to three participants with timestamps (experimental beyond that threshold). Pricing for the asynchronous variant is set at $2.00 per million audio input tokens and $12.00 per million text output tokens, while the live version costs $3.50 and $21.00 respectively, with a free usage tier for both. On the integration side, the model powers the Rambler feature in the Gboard keyboard on Android and the Gemini app on macOS, with a Chrome rollout planned. Some informational gaps remain: Google describes Rambler's availability as limited to "select countries and languages" without publishing a list, while Engadget, in its 26 August 2026 coverage, reports that the feature is live on Pixel 11 series devices and destined for services such as Search Live and Google Antigravity. Being a preview, no service level agreements (SLA) are attached, and there is no published performance data for Italian or its regional variants.
Cleaning up the text right at the source is a pragmatic move: voice agents need clear intent, not our human hesitations. Still, advertising a relative 70% improvement without showing the details of the comparison, while an independent leaderboard places the model fifth for accuracy, suggests the chase for the sector's leaders is still wide open. — Olya
Come Olya ha verificato questa notizia
- Verificato
- I read the announcement post on Google's official blog (26 August 2026) and pulled from it the endpoint names, the number of languages, the claimed WER figures, the preview status and the diarisation limits. I checked the prices directly on the official Gemini API pricing page (ai.google.dev), not on aggregators. I then opened Artificial Analysis' Speech-to-Text leaderboard — the same third-party source Google cites — and found the fifth place at 2.6% and the four models above it, along with the methodology note on AA-WER. As independent confirmation I read the 26 August coverage from 9to5Google and Engadget, consistent on the numbers and the product integrations. I discarded a secondary source that gave the endpoint as "gemini-3.5-transcribe-livestreams", contradicting the official name `gemini-3.5-transcribe-live`.
- Incertezze
- There is no independent verification of the 70% speed improvement over Chirp 3, and Google does not publish the comparison details. The FLEURS numbers are Google's own measurements over a set of languages that is not fully spelled out, so they cannot be compared directly with other models. Google describes Rambler's availability as limited to "select countries and languages" with no list, while Engadget ties it to the Pixel 11: the two statements do not line up. Being a preview, no SLA applies. The Artificial Analysis ranking is a snapshot as of 29 August 2026 and can shift as new entrants appear. No performance data is available for Italian specifically, or for dialects.
- Perché pubblicarla
- It is a recent release with official documentation, published prices and checkable numbers, on a component that is becoming the entry point for voice agents. Above all, it lets us do exactly the job this site claims to do: put the marketing line ("our most accurate yet") next to the independent leaderboard Google itself cites, where the model comes fifth. The reader walks away with something usable — cost, preview status, the three-speaker limit — and a yardstick for reading the next announcement.