Key takeaways
- Google announced EmbeddingGemma 2 on October 6, 2026. The official google/embeddinggemma-2 weights are available now on Hugging Face and Kaggle under Apache 2.0; Gemini Enterprise Agent Platform Model Garden support was described only as coming soon.
- The full model has 740M parameters: a 270M text path plus independently loadable 170M vision and 300M audio encoders. All supported modalities map into one native 768-dimensional space.
- Every input shares one 8,192-token budget. Google lists approximate single-modality maxima of 29 default-budget images, 58 video frames or 327 seconds of audio, so “8K context” should not be read as 8K text tokens plus unlimited media.
- Google reports MTEB Code rising from 68.76 for EmbeddingGemma 1 to 78.68 for version 2. These are vendor-run model-card results, not independent evidence or a guarantee for a private corpus.
- There is no per-token API price for downloaded open weights. Buyers still need to measure device compatibility, indexing time, storage, latency, energy and quality at the selected modality set, precision and vector dimension.
The weights are available now, while one managed channel is not
Google DeepMind launched EmbeddingGemma 2 on October 6, 2026 as an open multimodal embedding model built on Gemma 4 technology. The canonical model ID is google/embeddinggemma-2. Google linked live downloads on Hugging Face and Kaggle, and the reviewed Hugging Face repository identifies Google as publisher and Apache 2.0 as the license.
Google also named a future distribution channel: Gemini Enterprise Agent Platform Model Garden availability is “coming soon.” Do not put that promised route in a current architecture diagram. Today’s verifiable path is downloadable weights and local or self-managed inference through supported runtimes; access, billing and service behavior for the future managed listing remain unverified.
One vector space covers five input types
EmbeddingGemma 2 maps text, code, images, video, audio and interleaved combinations into a shared 768-dimensional vector space. That makes cross-modal retrieval possible without maintaining a separate embedding model and translation layer for every media type: a text or voice query can be compared with indexed images or video segments produced by the same model.
The architecture is modular. Google documents 270M parameters for text and code, 440M with vision, 570M with audio and 740M with both media encoders. Selective loading is a runtime contract, not model compression: omit an unused encoder in the library configuration so its weights are not loaded, then verify the resulting package and peak memory on the actual target.
The 8K context is shared across media
The context window is 8,192 tokens, four times the launch post’s stated context for EmbeddingGemma 1. Media consumes that same budget. The model card assigns 280 tokens to a default-budget image, 140 to a default video frame and 25 per audio second. Its approximate single-modality maxima are 29 images, 58 frames or 327 seconds of audio; interleaving text and media reduces each allowance.
Video defaults to one sampled frame per second and audio should be mono at 16 kHz. Vision can trade quality, latency and context consumption by changing its soft-token budget. Teams should therefore record preprocessing, sampling rate, segment boundaries and token budgets alongside the model ID. “Supports video” does not mean it semantically observes every frame of an arbitrary-length file.
Google reports stronger code retrieval, but no independent result yet
In Google’s full-precision evaluation, EmbeddingGemma 2 scores 78.68 on MTEB Code versus 68.76 for EmbeddingGemma 1, a 9.92-point absolute increase. The same table reports 61.36 on multilingual MTEB, 64.64 on MIEB Lite, 59.01 across MMEB v2, 69.54 on MSEB Retrieval and 49.39 on MAEB. These are vendor-reported results across different benchmark metrics; they should not be combined into one universal quality score.
The useful question is whether the model improves retrieval on the organization’s own corpus, queries and relevance judgments. Evaluate text, code and each media route separately; include multilingual and safety slices; compare end-to-end retrieval after chunking and preprocessing; and measure recall and ranking quality before a generator can hide weak retrieval behind fluent output.
Vector truncation saves storage at a measurable quality cost
Matryoshka Representation Learning lets deployments shorten the native 768-dimensional vector to 512, 256 or 128 dimensions. Google’s table reports multilingual MTEB moving from 61.36 at 768 dimensions to 60.41 at 256 and 57.89 at 128. Its MMEB v2 overall result moves from 59.01 to 56.24 and then 45.65. The sixfold storage claim applies to 128 versus 768 dimensions, not to free model inference.
Queries and corpus vectors must use the same dimension, and shortened vectors must be L2-normalized again. Google warns that skipping renormalization can silently degrade ranking rather than throw an error. Treat dimension as indexed-data schema: version it, test it and rebuild deliberately instead of changing it in place.
Precision and prompts are part of the model contract
Google says to run the model in bfloat16 or float32 and explicitly warns against float16 because the activation range can produce NaNs or silently degraded embeddings. That is a release-critical constraint for mobile converters and GPU pipelines. Add finite-value checks, norm checks and retrieval canaries after every export, quantization or runtime change.
Text tasks also use instruction prefixes. Retrieval uses different query and document forms, while classification, clustering and similarity use symmetric prompts. Images, audio and video do not take those text prefixes. Omitting the intended prompt still produces an embedding but may reduce precision, so prompts belong in the versioned indexing recipe rather than application folklore.
What builders should do next
Pilot EmbeddingGemma 2 when offline or private cross-modal retrieval is valuable and the target hardware can host the required encoders. Start with text-only or one media encoder, 256- and 768-dimensional indexes, bfloat16 where supported, and a labeled corpus. Record download commit or file hashes, runtime versions, prompts, preprocessing, vector dimension and normalization. Compare quality, cold start, steady-state latency, peak memory, index size and energy against the current retrieval stack.
Wait when a managed Model Garden route is mandatory, the device cannot meet measured memory or latency targets, or no representative multimodal evaluation set exists. Reject a rollout that treats vendor benchmarks as private-corpus proof, sends mismatched dimensions to one index, uses float16 without strict validation, or equates local processing with complete privacy: application logs, vector stores, caches and synced files still define the data boundary.
Copy-ready EmbeddingGemma 2 pilot record
Complete one record per device class, modality configuration and index version before production use.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Pilot, adopt, constrain, wait or reject; retrieval task, users, corpus, owner and review date.
google/embeddinggemma-2 revision, file hashes, Apache 2.0 review, runtime and conversion or quantization steps.
270M text, 440M text plus vision, 570M text plus audio, or full 740M; evidence unused encoders are omitted.
Task prefixes, chunking, title format, image token budget, video FPS, audio sample rate, segment length and interleaving rules.
768, 512, 256 or 128 dimensions; L2 normalization, index version, migration and rebuild procedure.
bfloat16 or float32 choice, finite-value and norm checks, unsupported float16 prevention, fallback and rollback.
Private labeled queries, per-modality recall and ranking metrics, languages, safety slices, baseline and acceptance thresholds.
Cold start, p50/p95 latency, peak and steady memory, index size, power, thermal behavior and offline test.
Local files, logs, caches, vector database, backups, synchronization, retention, deletion and access controls.
Test report, approvers, rollout cohort, monitoring, drift checks, incident owner and revalidation triggers.
Primary sources
Browse the publication-wide evidence index →
- EmbeddingGemma 2: an open, lightweight multimodal embedding modelGoogle DeepMind · Reviewed: October 6, 2026 launch date; architecture and parameter counts; modality, context and vector-size claims; Pixel memory figures; benchmark summary; availability and supported tooling · Retrieved · Supports: Google launched EmbeddingGemma 2 as a 740M-parameter Apache 2.0 model for text, code, image, video and audio embeddings, with downloadable weights and a modular 270M text-only configuration.
- EmbeddingGemma 2 model cardGoogle AI for Developers · Reviewed: Model overview; full benchmark and vector-truncation tables; task prefixes; selective encoder loading; numerical precision warning; multimodal context accounting; training data; safety; intended uses and limitations · Retrieved · Supports: The model card documents the 8,192-token shared context, 768-dimensional native output, supported 512/256/128 truncation, vendor-run evaluation scores, input limits, bfloat16 or float32 requirement and downstream safety responsibilities.
- EmbeddingGemma 2: The Developer GuideGoogle Developers Blog · Reviewed: October 6, 2026 publication date; modular architecture; Sentence Transformers 6.1 requirement; selective-loading examples; task prompts; multimodal inputs and truncation guidance · Retrieved · Supports: Google provides the model ID google/embeddinggemma-2, Sentence Transformers loading examples and effective parameter counts for text-only, text-plus-vision, text-plus-audio and full-multimodal configurations.
- google/embeddinggemma-2Hugging Face · Reviewed: Live repository identity and availability; publisher namespace; Apache 2.0 license label; Transformers and Sentence Transformers tags; model card and files navigation · Retrieved · Supports: The official Google repository was live under google/embeddinggemma-2 when reviewed and labeled for Apache 2.0, feature extraction, Transformers and Sentence Transformers.
Limitations
AccessAllGPT used public Google release material and the live official Hugging Face repository. We did not download or inspect weight files; run EmbeddingGemma 2; reproduce Google’s benchmark, memory, language or storage claims; compare it with independent models; test any modality, prefix, precision, truncation, quantization, converter, runtime or device; inspect the training corpus; verify filtering or bias; conduct a security, privacy or license review; calculate hardware or operating costs; or access the future Gemini Enterprise Agent Platform Model Garden listing. Benchmark names use different tasks and metrics and are not directly interchangeable. Repository files, integrations, license terms, managed availability and documentation can change.
Disclosures
AccessAllGPT received no Google account, hardware, weights, credits, early access, briefing, demo, benchmark artifact, review or compensation for this article. Google did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Google, Google DeepMind, Hugging Face or the benchmark projects cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- Google Launches Gemma 4 Argon With an Open-Weight Promise Still Pending
- Google Launches Gemini Guided Vision for Real-Time Camera Help
- Choose the Right RAG Chunk Overlap With Evidence
- Build an LLM Evaluation Platform You Can Move
- AI API Data Retention and Residency: Set the Procurement Gates
- AccessAllGPT Research methodology
- Evidence standards
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.