Gemini Embedding 2 logo on a black background with a wave of colored dots

Gemini Embedding 2 is a multilingual embedding model from Google. It encodes text, images, video, audio, and PDFs into one shared vector space. Because every input type lands in the same space, a text query can retrieve an image, and an image can retrieve a passage of audio, with no alignment step between modality-specific models. Google calls it an omni embedding model and distributes it through the Gemini API.

The design removes a common piece of retrieval plumbing. A conventional multimodal pipeline runs a text encoder, an image encoder, and sometimes an audio or video encoder. It maintains a separate index for each and reconciles the results after retrieval. Gemini Embedding 2 replaces that with one model and one index. Similarity scores become comparable across content types, so one threshold covers the whole corpus.

Retrieval no longer depends on how well the corpus was tagged. A query for “process documentation” can return a workflow diagram that nobody labeled with those words, because the match is on the content of the image. That matters most in media libraries, e-commerce catalogs, and document sets that accumulated inconsistent tagging over years.

Inputs and limits

Five rows listing the limits per input type: text, images, video, audio and PDFs.

Each modality has its own limit:

  • Text: 8,192 tokens, four times the limit of earlier Gemini embedding models.
  • Images: six per request, PNG or JPEG.
  • Video: 120 seconds, MP4 or MOV.
  • Audio: embedded directly, with no transcription step. Supply it as inline data or through the Files API.
  • PDFs: six pages.

The model embeds a PDF as a unit. A page that combines prose, a chart, and a table keeps the relationships among those elements, since layout is part of what the model reads. Extracting the text first would discard that layout.

Dimensions

Three nested boxes of 3,072, 1,536 and 768 dimensions, showing how a shortened vector keeps its core information at lower cost.

The default output is 3,072 dimensions. The output_dimensionality parameter lowers that, and 1,536 and 768 are the common alternatives. Truncation works because the model uses Matryoshka Representation Learning. This method concentrates information in the early dimensions of the vector, so a shortened vector still supports similarity comparisons. Smaller vectors cost less to store and search, but they give up some precision.

Calling the model

A developer at a standing desk with two monitors, working through an embedding call and checking the results.

Embeddings come from the embed_content method on the Gemini API. Client libraries exist for Python, JavaScript, and Go. A typical Python call reads binary content from disk, sends it, receives a 3,072-dimension vector, and normalizes the vector before any similarity computation.

The SDK is split into submodules. models handles generation calls such as generateContent and generateImages, caches handles reusable prompts, chats handles multi-turn state, and files handles uploading and referencing assets. Errors surface through an ApiError class. You instantiate clients through the GoogleGenAI class and set the API version, v1 or v1alpha, at that point. Authentication uses an API key. You can enter the key once and store it instead of passing it through environment variables on every run.

You can reach the model through Google AI Studio and Vertex AI as well as the Gemini Developer API. Google publishes stable, preview, latest, and experimental versions, so production code can pin a version while testing tracks a newer one. Teams that want to avoid managing a backend can use MindStudio, a no-code environment that wraps Gemini models and handles the API connections and authentication.

Vectors go into whatever vector database you already use. Pinecone, Weaviate, Qdrant, ChromaDB, and pgvector all appear in the documentation and in community write-ups. The Batch API processes large jobs at higher throughput and lower cost per embedding, which matters when the corpus is a video archive instead of a few thousand text documents.

Retrieval-augmented generation

Text, charts and diagrams in one index, where a single query retrieves them together for the generation step.

The single index simplifies retrieval-augmented generation (RAG). One query retrieves relevant chunks across every content type, so the generation step receives charts and diagrams along with the text around them. The pipeline needs no post-retrieval fusion logic to merge results from separate indexes.

For technical documentation, research papers, and financial reports, this changes what the retrieval step can find. A figure that carries the result of an experiment becomes retrievable directly as well as through its caption. Legal discovery is the case Google cites most often: scanned documents, image attachments, and recordings sit in one index, and one query covers all of them.

Representations stay consistent across languages, so cross-lingual retrieval works without a translation step in front of the query.

Limitations

Five caution rows listing the limits of Gemini Embedding 2: no tuning, chunked audio, downsampled images, weaker video and cost.

Model tuning is not available for Gemini models. Google lists it as in progress, so architectures must work with the pretrained model as published.

Resolution and duration limit non-text inputs. You must chunk long audio, and high-resolution images are downsampled before embedding. Video embeddings do not match models built specifically for video understanding: action recognition in longer clips and fine-grained audio classification are both weaker.

Embedding a large mixed-media corpus costs money and runs against rate limits. Plan for both before you start the first bulk job.

Safety settings and SafetyRating categories apply across the Gemini API. Google documents the permitted content, prompts, and model behavior separately from the embedding model itself.

Some of what determines reliability in production sits outside the model. Session management, retry behavior, HTTPS enforcement, session ID rotation on model fallback, and the ContextManager in Gemini’s CLI tooling all affect whether an embedding pipeline stays up. End-to-end reliability depends on the maturity of the whole stack.

Release status

A project manager planning around a wall calendar, showing the need to confirm release status before committing to a schedule.

Experimental embedding releases appeared in late 2024, followed by public preview, with general availability discussed for 2026. The public record is inconsistent on this point. Some accounts state that Google has made no official release announcement and that availability during certain phases carried no guarantee. Confirm the current status against Google’s own release notes before you commit to a deployment schedule.

Share This
AI & Search SystemsEmbeddingsGemini Embedding 2