EmbeddingGemma 2 Studio

Google DeepMind • Released Oct 2026 • Apache 2.0

EmbeddingGemma 2 Multimodal Search & MRL Playground

Test unified text, image & audio embeddings in your browser with WebGPU. Experiment with Matryoshka dimension truncation (128d–768d) and compute cross-modal cosine similarity with zero server latency.

🧬 740M Params (270M Text + 170M Vision + 300M Audio) šŸ“ 768-Dim Unified Embedding Space ⚔ 83% RAM Saved via Matryoshka 128d MRL šŸ”’ 100% Client-Side Privacy

Live Multimodal Similarity Sandbox WebGPU Ready

Select a real-world scenario preset or drag & drop your own files to evaluate cross-modal vector distances.

Query Modality: Natural Language / Code Tokens: ~12
VS
Cosine Angle
Vision (Image)
Dog on beach

Click or drag & drop to replace with your image / audio file

Subject: Golden Retriever playing on sunny beach Reset Preset
Cosine Similarity
0.884 / 1.000

✨ High Semantic Match (High Confidence)

Vector Alignment: 27.8° Euclidean Distance: 0.481
Inference: WebGPU fp16 (38ms) • Normalized: L2 (||v|| = 1.0)

Matryoshka Truncation Slider (MRL) MTEB Feature

Truncate dimensions on-the-fly without model fine-tuning. Observe memory savings vs similarity impact.

Current Dimension: 768d (Full)
128d (Ultra-Fast) 256d (Compact) 512d (High Fidelity) 768d (Full Dimension)
Vector Size
3.07 KB / vec
1M Vectors RAM
3.07 GB
Storage Reduction
0.0% (Baseline)
NDCG@10 Retention
100.0%

Production Integration Snippets

Copy ready-to-run implementation code for Python, Transformers.js (WebGPU), Ollama, and cURL.

from sentence_transformers import SentenceTransformer
from PIL import Image
import numpy as np

# 1. Load Google DeepMind EmbeddingGemma 2 (740M modular parameters)
model = SentenceTransformer("google/embeddinggemma-2-740m", trust_remote_code=True)

# 2. Asymmetric retrieval: Prefix search queries with 'task: search_query | '
query_text = "task: search_query | a golden retriever playing on sunny beach"
query_embedding = model.encode(query_text)

# 3. Multimodal ingestion: Target image is mapped into identical 768d vector space
target_image = Image.open("sample_photo.jpg")
image_embedding = model.encode(target_image)

# 4. Optional Matryoshka Representation Learning (MRL) truncation to 256 dimensions
query_256 = query_embedding[:256]
image_256 = image_embedding[:256]

# 5. Compute cosine similarity
query_norm = query_256 / np.linalg.norm(query_256)
image_norm = image_256 / np.linalg.norm(image_256)
similarity = float(np.dot(query_norm, image_norm))

print(f"Cosine Similarity (256d MRL): {similarity:.4f}")

Multimodal Embedding Models Comparison (2026)

How EmbeddingGemma 2 compares against OpenAI CLIP, SigLIP, and Voyage AI on MTEB benchmarks and edge deployment efficiency.

Model Name Params Dimensions Supported Modalities MRL Truncation On-Device WebGPU License
EmbeddingGemma 2 740M (Modular) 768d (128d-768d) Text, Code, Img, Audio, Video āœ… Native (96.8% retain) āœ… Supported Apache 2.0
OpenAI CLIP ViT-L/14 428M 768d Text, Image āŒ No āš ļø Heavy MIT
Google SigLIP-SO400M 400M 1152d Text, Image āŒ No Supported Apache 2.0
Voyage Multimodal 3 Proprietary 1024d Text, Image, PDF āœ… Supported āŒ Cloud API Only Commercial API
Nomic Embed Vision v1.5 93M 768d (64d-768d) Text, Image āœ… Supported āœ… Supported Apache 2.0

Frequently Asked Questions

What makes EmbeddingGemma 2 unique compared to other embedding models? ā–¼

Traditional embedding models are almost exclusively unimodal (only text, e.g. text-embedding-3 or BGE-M3) or dual-modal (CLIP/SigLIP for text and images only). EmbeddingGemma 2 features a truly unified space spanning text, source code, images, audio clips, and video keyframes into a single 768-dimensional vector space. Furthermore, its modular encoder architecture allows loading only the modalities required on lightweight edge devices.

How much RAM and storage can I save using Matryoshka (MRL) truncation? ā–¼

By default, a 768-dimensional FP32 vector occupies ~3.07 KB. Storing 1,000,000 vectors requires roughly 3.07 GB in RAM or Vector DB index (Qdrant / Milvus / Pinecone). By truncating to 128 dimensions, each vector consumes only ~0.51 KB (or 0.25 KB in FP16), saving up to 83.3% of vector index storage with minimal precision degradation (>96.8% NDCG retention).

Why do I need to prepend 'task: search_query | ' to queries? ā–¼

EmbeddingGemma 2 is optimized for asymmetric semantic search. Search queries are typically short, exploratory phrases, while targets are complete images, paragraphs, or audio recordings. Prepending the required prompt prefix (e.g. task: search_query | or task: classification | ) instructs the attention layers to compute representations tailored specifically for query-to-document alignment.

Can this run completely offline without an internet connection? ā–¼

Yes! Once model weights are fetched via Transformers.js or downloaded in Ollama / LiteRT, inference runs entirely offline on your local GPU / NPU. No text or image files are ever uploaded to cloud servers, making it ideal for strict enterprise privacy, HIPAA, or on-device personal assistant apps.

Share This Playground with AI Engineers

Found this demo helpful? Share your test results to X/Twitter, Hacker News, or Reddit.