Embedding Gemma 2: On-Device Multimodal Embeddings and Retrieval Demonstrations

1 件の動画 · 更新: 3時間前
Introducing EmbeddingGemma 2: An open model for natively multimodal embeddings 📺 Introducing EmbeddingGemma 2: An open model for natively multimodal embeddings ⏱ 3:09📅 2026/10/06 16:10🌐 English Translation

Embedding Gemma 2: On-Device Multimodal Embeddings and Retrieval Demonstrations

▼

This video introduces Embedding Gemma 2, an open lightweight multimodal embedding model that maps text, images, video, and audio into a unified embedding space. It explains the model design and demonstrates on-device retrieval workflows using Google AI Edge tools.

Model design and capabilities:
- Unified embedding space for text, images, video, audio, or mixtures
- Under 1B parameters, up to 740M modular design, 768-dimensional vectors, Matryoshka reduction to 128 dimensions, 8K context

Demos and on-device workflows:
- Google AI Edge Gallery for natural-language media search
- Video Moment Finder for locating moments within video
- AI Edge Foresight for audio-stream embedding, question detection, and local RAG with models like Gemma 4

Tuning and access:
- Domain-specific fine-tuning for legal contracts, medical imaging, or tech product catalogs
- Download model weights from Hugging Face and explore Gemma guide notebooks

The video is suitable for developers and ML engineers interested in on-device multimodal retrieval, privacy-preserving RAG, and compact embedding models. It provides an overview of the model's intended use cases and where to access its weights and examples.

この動画を紹介した Google for Developers の最新動画も、紹介付きで読めます。

📄 このページの紹介文は AI が独自に生成したものであり、著作権をはじめとする他者の権利(商標権・名誉権・プライバシー等)を侵害しないよう配慮しています。動画の著作権は各作成者に帰属します。