EmbeddingGemma 2 Brings Private Multimodal Search On-Device

Google’s 740M-parameter EmbeddingGemma 2 runs local search across text, code, images, audio and video. Here are the memory, vector-size and deployment tradeoffs.
EmbeddingGemma 2 banner showing text, image, audio and video inputs connected in a shared embedding space
EmbeddingGemma 2 maps text, code, images, audio and video into a shared vector space. Image: Google DeepMind.

Google DeepMind released EmbeddingGemma 2 on October 6, an open 740-million-parameter model designed to search across text, code, images, audio and video without sending the source material to a cloud service. The model maps every supported input into the same 768-dimensional vector space, so a text query can retrieve a video frame, a voice recording can locate an image, and an app can search several media types with one index.

The release is available under Apache 2.0 through Hugging Face and is already supported by Transformers and SentenceTransformers. Google is positioning it for local semantic search, retrieval-augmented generation, classification and routing on phones and laptops. Its real appeal is architectural: developers no longer need to chain speech transcription, image captioning and a text-only embedding model just to make mixed media searchable.

What EmbeddingGemma 2 actually does

An embedding model does not answer a question or generate prose. It converts an input into a list of numbers that represents its meaning. Items with related meaning should land near one another in that mathematical space, allowing a vector database or a simple similarity calculation to rank the closest matches.

EmbeddingGemma 2 uses a 270M-parameter text component, a 170M vision encoder and a 300M audio encoder. Those modules can be loaded separately. A text-only deployment therefore uses 270M parameters; text plus images uses 440M; text plus audio uses 570M; and the full configuration uses all 740M. Video is sampled as frames by the vision encoder, while its audio track can be represented through the audio tower.

All modalities share an 8,192-token context window. At Google’s default settings, that budget holds roughly 29 images, 58 video frames or 327 seconds of audio when no other input is present. Mixed inputs consume the same shared allowance. Developers can also interleave media and text in one representation, such as a product description followed by photos and a short demonstration video.

The local memory claim needs context

Google reports about 191MB of active RAM for quantized text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Those figures describe model weights in Google’s test configuration, not the complete memory requirement of an application. The runtime, media buffers, tokenizer, vector index and the host app all add overhead.

Even so, the footprint makes genuinely local uses plausible: searching a personal photo library without uploading it, locating moments inside recorded meetings, indexing a private codebase, or routing an offline voice command by comparing it with a set of action descriptions. Local processing can reduce latency and exposure of raw data, although it does not automatically make an app private; telemetry, crash reports, synchronization and the storage of embeddings still require their own review.

Vector size is a practical cost control

The model produces 768-number vectors by default, but it was trained with Matryoshka Representation Learning, which lets developers keep the leading 512, 256 or 128 dimensions. Shorter vectors take less storage and make similarity searches cheaper.

Google’s model card shows the central tradeoff. At 256 dimensions, vector storage falls to one-third of the full size while the multilingual MTEB score moves from 61.36 to 60.41 and the code score from 78.68 to 76.18. At 128 dimensions, storage falls to one-sixth, but the overall multimodal score drops from 59.01 to 45.65. That makes 256 dimensions a sensible starting point for many mixed-media tests; 128 dimensions needs workload-specific validation and is better suited to compact text indexes.

There is an important implementation trap. A shortened vector must be L2-normalized again after truncation. Simply slicing a 768-dimensional vector produces plausible-looking similarity scores that are wrong rather than triggering an obvious error. Queries and indexed documents must also use the same dimension.

Code retrieval is the largest measured improvement

Google reports an MTEB Code score of 78.68, up from 68.76 for the first EmbeddingGemma. Its multilingual text score barely changed, moving from 61.15 to 61.36. The bigger story is therefore not a broad leap in text retrieval; it is the addition of native media encoders and a 9.92-point improvement in Google’s reported code benchmark.

The full-precision checkpoint also scored 64.64 on MIEB Lite for images, 50.67 on the video portion of MMEB v2, and 69.54 on the MSEB retrieval benchmark for audio. These are vendor-reported results across different datasets and metrics, not directly comparable percentages. Teams should treat them as screening evidence and rerun retrieval tests on their own documents, languages, image styles and audio conditions.

Two configuration choices can quietly damage results

Google’s documentation identifies two failure modes that deserve attention before a production trial.

  • Do not run the model in float16. Its activation range can exceed float16, leading to NaN values or silently degraded embeddings. Google recommends bfloat16 on compatible hardware and float32 on most CPUs.
  • Use the task prefixes. Search queries and indexed documents are trained with different instructions. SentenceTransformers can apply names such as SearchQuery, Document and CodeRetrieval. Omitting the prefix still returns a vector, but lowers retrieval precision.
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "google/embeddinggemma-2",
    config_kwargs={
        "vision_config": None,
        "audio_config": None,
    },
)

query = model.encode(
    "Where is authentication initialized?",
    prompt_name="CodeRetrieval",
    truncate_dim=256,
    normalize_embeddings=True,
)

document = model.encode(
    "title: auth.py | text: def create_session(...):",
    truncate_dim=256,
    normalize_embeddings=True,
)

This text-only example deliberately disables the unused media encoders and keeps 256 dimensions. A real evaluation should compare the model with the incumbent system on recall at a fixed result count, latency on target hardware, index size, power use and performance on hard negatives, not just inspect a handful of attractive matches.

Where it fits and where it does not

EmbeddingGemma 2 is a retrieval component, not a complete RAG stack. It does not chunk files, maintain access controls, update an index, filter unsafe results or generate the final answer. Applications still need a document pipeline, vector store, permission-aware retrieval and, when answers are required, a generative model.

It is strongest where one compact local model can replace several preprocessing services. A media app could let someone speak a description to find a moment in a video. A field tool could match a photo and note against an offline repair manual. A private knowledge app could search code, screenshots and recorded explanations together. These are more meaningful tests than treating the release as a drop-in upgrade for every text-search system.

Google has also published mobile demonstrations through AI Edge Gallery and Mac examples through AI Edge Foresight. The developer launch guide describes semantic search, visual keyframe retrieval and condition-trigger workflows. Because the weights and implementation are public, teams can benchmark those patterns without committing data to a hosted embedding API.

The release makes multimodal search much easier to prototype on ordinary hardware. Whether it replaces a specialized text, vision or audio model will depend on the workload. The first tests should be small and measurable: choose a representative corpus, decide which modalities are truly needed, compare 768 and 256 dimensions, avoid float16, and inspect retrieval errors before building the rest of the product around the index.

Previous Post
Constellation Energy's Braidwood and Byron nuclear power plants in Illinois

Google’s Constellation Nuclear Deal: What the 3.59 GW Figure Really Means

Related Posts