📺 Introducing EmbeddingGemma 2: An open model for natively multimodal embeddings
This video introduces Embedding Gemma 2, an open lightweight multimodal embedding model that maps text, images, video, and audio into a unified embedding space. It explains the model design and demonstrates on-device retrieval workflows using Google AI Edge tools.
Model design and capabilities:
- Unified embedding space for text, images, video, audio, or mixtures
- Under 1B parameters, up to 740M modular design, 768-dimensional vectors, Matryoshka reduction to 128 dimensions, 8K context
Demos and on-device workflows:
- Google AI Edge Gallery for natural-language media search
- Video Moment Finder for locating moments within video
- AI Edge Foresight for audio-stream embedding, question detection, and local RAG with models like Gemma 4
Tuning and access:
- Domain-specific fine-tuning for legal contracts, medical imaging, or tech product catalogs
- Download model weights from Hugging Face and explore Gemma guide notebooks
The video is suitable for developers and ML engineers interested in on-device multimodal retrieval, privacy-preserving RAG, and compact embedding models. It provides an overview of the model's intended use cases and where to access its weights and examples.
この動画を紹介した Google for Developers の最新動画も、紹介付きで読めます。
📄 このページの紹介文は AI が独自に生成したものであり、著作権をはじめとする他者の権利(商標権・名誉権・プライバシー等)を侵害しないよう配慮しています。動画の著作権は各作成者に帰属します。