Google Unveils Gemini 3.5 Transcribe for Accurate Text Conversion in Noisy Environments
Google has officially launched Gemini 3.5 Transcribe, the latest addition to its Gemini series, positioning it as the most accurate speech recognition model in the lineup. Now available to developers, enterprises, and general users, the model addresses persistent industry challenges, including background noise, specialized terminology, and natural speech patterns.
Featuring robust generalization capabilities, Gemini 3.5 Transcribe excels in filtering background noise and handling complex vocabulary across specialized fields. The model has already enhanced performance for the Gemini app on macOS and the Rambler feature in Gboard for Android. Developers can leverage this tool to build innovative solutions, such as voice-enabled agents, real-time captioning systems, and post-call analysis workflows.

Optimized for the nuances of daily conversation, the model accurately captures semantic intent and identifies custom vocabulary to assist users in completing tasks via voice. It includes practical features like automatic self-correction, intelligent removal of filler words (e.g., "um," "uh"), and automatic text formatting. Furthermore, it can delegate complex tasks, such as file analysis and image generation, to other Gemini models, offering an end-to-end multi-model collaboration experience.
Google’s public data indicates that Gemini 3.5 Transcribe achieves an average word error rate (WER) of just 4.0% in streaming scenarios and 2.6% in non-streaming scenarios. It maintains high recognition stability even in noisy environments, delivering superior accuracy for critical alphanumeric information like order numbers and postal codes.
The model also excels in multilingual support and personalization, supporting 85 languages with adaptation to various regional accents and dialects. For preprocessed audio, it offers speaker identification with timestamps for up to three speakers. Additionally, it fully supports user-defined vocabulary lists, effectively handling professional terminology, unique spellings, and industry-specific names.
Regarding ecosystem integration, developers can access Gemini 3.5 Transcribe in preview via the Gemini API on Google AI Studio and Google Antigravity, while general users can try it on Android and macOS. Google plans to integrate the model into Chrome, enabling voice input in any web field. This rollout follows the launch of Gemini 3.7 Flash, the expansion of Gemini to all U.S. Android users, and its integration into Waymo autonomous vehicles, reinforcing Google’s rapid advancement of its full-scenario AI product matrix.
Related article
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Related Special Topic Recommendations
Comments (0)
0/500
Google has officially launched Gemini 3.5 Transcribe, the latest addition to its Gemini series, positioning it as the most accurate speech recognition model in the lineup. Now available to developers, enterprises, and general users, the model addresses persistent industry challenges, including background noise, specialized terminology, and natural speech patterns.
Featuring robust generalization capabilities, Gemini 3.5 Transcribe excels in filtering background noise and handling complex vocabulary across specialized fields. The model has already enhanced performance for the Gemini app on macOS and the Rambler feature in Gboard for Android. Developers can leverage this tool to build innovative solutions, such as voice-enabled agents, real-time captioning systems, and post-call analysis workflows.

Optimized for the nuances of daily conversation, the model accurately captures semantic intent and identifies custom vocabulary to assist users in completing tasks via voice. It includes practical features like automatic self-correction, intelligent removal of filler words (e.g., "um," "uh"), and automatic text formatting. Furthermore, it can delegate complex tasks, such as file analysis and image generation, to other Gemini models, offering an end-to-end multi-model collaboration experience.
Google’s public data indicates that Gemini 3.5 Transcribe achieves an average word error rate (WER) of just 4.0% in streaming scenarios and 2.6% in non-streaming scenarios. It maintains high recognition stability even in noisy environments, delivering superior accuracy for critical alphanumeric information like order numbers and postal codes.
The model also excels in multilingual support and personalization, supporting 85 languages with adaptation to various regional accents and dialects. For preprocessed audio, it offers speaker identification with timestamps for up to three speakers. Additionally, it fully supports user-defined vocabulary lists, effectively handling professional terminology, unique spellings, and industry-specific names.
Regarding ecosystem integration, developers can access Gemini 3.5 Transcribe in preview via the Gemini API on Google AI Studio and Google Antigravity, while general users can try it on Android and macOS. Google plans to integrate the model into Chrome, enabling voice input in any web field. This rollout follows the launch of Gemini 3.7 Flash, the expansion of Gemini to all U.S. Android users, and its integration into Waymo autonomous vehicles, reinforcing Google’s rapid advancement of its full-scenario AI product matrix.
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation





Home






