📺 Build voice-first apps with Gemini 3.5 Transcribe
This content introduces the launch of a new LLM-based transcription model powered by Gemini, available on both the Interactions API and the Live API. It highlights enhanced accuracy for complex alphanumeric data and robust multilingual support.
- Model Capabilities: Utilizes deep audio understanding and reasoning to accurately transcribe challenging inputs like email addresses and phone numbers.
- Multilingual Support: Recognizes and transcribes over 70 languages correctly, even when language hints are set to English.
- Availability: The model is accessible via the Interactions API and supports live transcription through the Live API.
Developers and engineers can leverage this model to improve transcription reliability for specific formats and diverse languages.
この動画を紹介した Google for Developers の最新動画も、紹介付きで読めます。
📄 このページの紹介文は AI が独自に生成したものであり、著作権をはじめとする他者の権利(商標権・名誉権・プライバシー等)を侵害しないよう配慮しています。動画の著作権は各作成者に帰属します。