Live Multimodal Similarity Sandbox WebGPU Ready
Select a real-world scenario preset or drag & drop your own files to evaluate cross-modal vector distances.
Click or drag & drop to replace with your image / audio file
⨠High Semantic Match (High Confidence)
Matryoshka Truncation Slider (MRL) MTEB Feature
Truncate dimensions on-the-fly without model fine-tuning. Observe memory savings vs similarity impact.
Production Integration Snippets
Copy ready-to-run implementation code for Python, Transformers.js (WebGPU), Ollama, and cURL.
from sentence_transformers import SentenceTransformer
from PIL import Image
import numpy as np
# 1. Load Google DeepMind EmbeddingGemma 2 (740M modular parameters)
model = SentenceTransformer("google/embeddinggemma-2-740m", trust_remote_code=True)
# 2. Asymmetric retrieval: Prefix search queries with 'task: search_query | '
query_text = "task: search_query | a golden retriever playing on sunny beach"
query_embedding = model.encode(query_text)
# 3. Multimodal ingestion: Target image is mapped into identical 768d vector space
target_image = Image.open("sample_photo.jpg")
image_embedding = model.encode(target_image)
# 4. Optional Matryoshka Representation Learning (MRL) truncation to 256 dimensions
query_256 = query_embedding[:256]
image_256 = image_embedding[:256]
# 5. Compute cosine similarity
query_norm = query_256 / np.linalg.norm(query_256)
image_norm = image_256 / np.linalg.norm(image_256)
similarity = float(np.dot(query_norm, image_norm))
print(f"Cosine Similarity (256d MRL): {similarity:.4f}")
import { pipeline, env } from '@huggingface/transformers';
// Enable WebGPU device and multi-threaded WebAssembly
env.backends.onnx.wasm.numThreads = 4;
// 1. Initialize feature-extraction pipeline with browser WebGPU
const embedder = await pipeline('feature-extraction', 'google/embeddinggemma-2-740m', {
device: 'webgpu',
dtype: 'fp16',
});
// 2. Compute text query embedding directly on user hardware (zero backend latency)
const queryOutput = await embedder('task: search_query | vintage brown leather jacket', {
pooling: 'mean',
normalize: true,
});
// 3. Extract 768-dimensional float array
const queryVector = Array.from(queryOutput.data);
// 4. Truncate with Matryoshka (e.g. 128 dimensions for hyper-fast vector cache)
const mrl128Vector = queryVector.slice(0, 128);
console.log('On-device 128d vector computed:', mrl128Vector.length);
# 1. Pull EmbeddingGemma 2 locally via Ollama
ollama pull embeddinggemma-2:740m
# 2. Run local embedding calculation via CLI
ollama run embeddinggemma-2 "task: search_query | high efficiency sorting algorithms"
# 3. Call local REST API from your backend app
curl http://localhost:11434/api/embeddings -d '{
"model": "embeddinggemma-2:740m",
"prompt": "task: search_query | microservices event driven architecture"
}'
# Hugging Face Inference Endpoints / OpenRouter API Call
curl https://api-inference.huggingface.co/models/google/embeddinggemma-2-740m \
-H "Authorization: Bearer $HF_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"inputs": {
"source_sentence": "task: search_query | cute golden retriever puppy on beach",
"sentences": [
"A playful dog near the ocean waves",
"A red sports car on race track"
]
}
}'
Multimodal Embedding Models Comparison (2026)
How EmbeddingGemma 2 compares against OpenAI CLIP, SigLIP, and Voyage AI on MTEB benchmarks and edge deployment efficiency.
| Model Name | Params | Dimensions | Supported Modalities | MRL Truncation | On-Device WebGPU | License |
|---|---|---|---|---|---|---|
| EmbeddingGemma 2 | 740M (Modular) | 768d (128d-768d) | Text, Code, Img, Audio, Video | ā Native (96.8% retain) | ā Supported | Apache 2.0 |
| OpenAI CLIP ViT-L/14 | 428M | 768d | Text, Image | ā No | ā ļø Heavy | MIT |
| Google SigLIP-SO400M | 400M | 1152d | Text, Image | ā No | Supported | Apache 2.0 |
| Voyage Multimodal 3 | Proprietary | 1024d | Text, Image, PDF | ā Supported | ā Cloud API Only | Commercial API |
| Nomic Embed Vision v1.5 | 93M | 768d (64d-768d) | Text, Image | ā Supported | ā Supported | Apache 2.0 |
Frequently Asked Questions
What makes EmbeddingGemma 2 unique compared to other embedding models? ā¼
Traditional embedding models are almost exclusively unimodal (only text, e.g. text-embedding-3 or BGE-M3) or dual-modal (CLIP/SigLIP for text and images only). EmbeddingGemma 2 features a truly unified space spanning text, source code, images, audio clips, and video keyframes into a single 768-dimensional vector space. Furthermore, its modular encoder architecture allows loading only the modalities required on lightweight edge devices.
How much RAM and storage can I save using Matryoshka (MRL) truncation? ā¼
By default, a 768-dimensional FP32 vector occupies ~3.07 KB. Storing 1,000,000 vectors requires roughly 3.07 GB in RAM or Vector DB index (Qdrant / Milvus / Pinecone). By truncating to 128 dimensions, each vector consumes only ~0.51 KB (or 0.25 KB in FP16), saving up to 83.3% of vector index storage with minimal precision degradation (>96.8% NDCG retention).
Why do I need to prepend 'task: search_query | ' to queries? ā¼
EmbeddingGemma 2 is optimized for asymmetric semantic search. Search queries are typically short, exploratory phrases, while targets are complete images, paragraphs, or audio recordings. Prepending the required prompt prefix (e.g. task: search_query | or task: classification | ) instructs the attention layers to compute representations tailored specifically for query-to-document alignment.
Can this run completely offline without an internet connection? ā¼
Yes! Once model weights are fetched via Transformers.js or downloaded in Ollama / LiteRT, inference runs entirely offline on your local GPU / NPU. No text or image files are ever uploaded to cloud servers, making it ideal for strict enterprise privacy, HIPAA, or on-device personal assistant apps.
Share This Playground with AI Engineers
Found this demo helpful? Share your test results to X/Twitter, Hacker News, or Reddit.