Google Unveils Gemini 3.5 Transcribe, Elevating Speech-to-Text Precision to New Heights
1 day ago / Read about 0 minute
Author:小编   

Google has recently rolled out a cutting-edge addition to its Gemini series—the Gemini 3.5 Transcribe, a speech-to-text model touted as the most accurate within the series. Tailored for developers, businesses, and everyday users alike, this model marks a significant leap forward in transcription capabilities, excelling in challenging environments such as noisy settings, complex technical jargon, and natural spoken language. Not only can it discern the speaker's intent and automatically structure the output, but it also adeptly identifies self-corrections, filters out filler words, and transforms spoken content into neatly formatted text.

Moreover, the model boasts support for over 85 languages, encompassing mixed-language speech scenarios, and empowers users to incorporate custom vocabulary to further refine recognition accuracy. Gemini 3.5 Transcribe has seamlessly been integrated into the Android iteration of Gboard and the Mac version of the Gemini app, while developers can also access a preview through the Gemini API. Plans are also underway to introduce this innovative technology to the Chrome browser, enabling web-based voice input capabilities.

  • C114 Communication Network
  • Communication Home
7 X 24 Track global technological trends
Hot Topic