Google Unveils Gemini Embedding2: Native Multimodal Model Unifies Semantic Spaces
Google has recently unveiled its new native multimodal embedding model, Gemini Embedding2. It can map text, images, videos, audio, and PDF documents into a shared semantic vector space, designed to streamline complex AI data workflows and improve multimodal retrieval and understanding. This represents a key advancement for Google in embedding technology, transitioning from single-modality text to unified multimodal semantic modeling.

Previously, in July 2025, Google introduced the gemini-embedding-001 text embedding model. It supported over 100 languages and achieved top results on the MTEB multilingual benchmark. The new Gemini Embedding2 builds on the Gemini architecture but significantly expands its scope. It now processes five different modalities—text, images, video, audio, and PDFs—and projects them into a single vector space. This allows for direct semantic comparisons across different types of media without needing multiple specialized models or extra processing steps. This capability is particularly valuable for applications like semantic search, retrieval-augmented generation (RAG), sentiment analysis, and data clustering.
Regarding input capabilities, the new model supports up to 8192 text tokens, quadruple the previous 2048-token limit. It can handle up to six PNG or JPEG images per request, videos up to 120 seconds long, and PDF documents up to six pages. A notable feature is Gemini Embedding2's native support for audio processing, eliminating the need for speech-to-text conversion and avoiding potential information loss from transcription. Google also introduced "interleaved input" technology, enabling developers to combine multiple modalities in one request—like mixing images with descriptive text—to better capture the semantic relationships between them.

Architecturally, the model continues to employ Matryoshka Representation Learning (MRL). This technique uses a hierarchical structure to dynamically adjust vector dimensions. The default embedding dimension is 3072, with optional configurations of 1536 and 768 available, giving developers flexibility to balance retrieval accuracy with storage efficiency.
Google's benchmark results indicate that Gemini Embedding2 delivers leading performance across text, image, video, and speech tasks. For instance, in text-video retrieval, it scores 68.8, outperforming Amazon Nova2Multimodal Embeddings (60.3) and Voyage Multimodal3.5 (55.2). In text-image comparison, it achieves a score of 93.4, significantly ahead of Amazon's model score of 84.0.
Gemini Embedding2 is currently accessible to developers via the Gemini API and Vertex AI. It integrates with popular frameworks and vector databases like LangChain, LlamaIndex, Haystack, Weaviate, Qdrant, ChromaDB, and Vector Search. To help developers get started, Google provides interactive Colab notebooks and lightweight multimodal semantic search demonstrations.

The competition in multimodal embedding is heating up. Notably, in late February of this year, the AI search engine Perplexity released its open-source embedding models, pplx-embed-v1 and pplx-embed-context-v1.
Related article
MiniMax Unveils 10x Team Program to Incentivize Global AI Experts
MiniMax (Xiyu Technology), the General Artificial Intelligence Lab, has officially launched "10x Team," a global talent collaboration initiative. This program aims to recruit top experts across industries to explore the deep application of large mode
South Korea Breaks Ground on National AI Computing Center, Investing 2.5 Trillion Won with 2028 Target
South Korean outlet EtNews reports that groundbreaking for the Korea AI Computing Center (KOACC) took place on August 3 at the Solar City data center park in Sunan, Jeollanam-do. Backed by a total investment of 2.5 trillion KRW (roughly 11.838 billio
Six Tech Giants Back Linux Foundation With $12.5M to Tackle AI Vulnerability Noise
To tackle the flood of low-quality security reports produced by AI automation tools, six major tech companies—Anthropic, Amazon (AWS), GitHub, Google, Microsoft, and OpenAI—have collectively contributed $12.5 million in funding to Linux Foundation in
Related Special Topic Recommendations
Comments (0)
0/500
Google has recently unveiled its new native multimodal embedding model, Gemini Embedding2. It can map text, images, videos, audio, and PDF documents into a shared semantic vector space, designed to streamline complex AI data workflows and improve multimodal retrieval and understanding. This represents a key advancement for Google in embedding technology, transitioning from single-modality text to unified multimodal semantic modeling.

Previously, in July 2025, Google introduced the gemini-embedding-001 text embedding model. It supported over 100 languages and achieved top results on the MTEB multilingual benchmark. The new Gemini Embedding2 builds on the Gemini architecture but significantly expands its scope. It now processes five different modalities—text, images, video, audio, and PDFs—and projects them into a single vector space. This allows for direct semantic comparisons across different types of media without needing multiple specialized models or extra processing steps. This capability is particularly valuable for applications like semantic search, retrieval-augmented generation (RAG), sentiment analysis, and data clustering.
Regarding input capabilities, the new model supports up to 8192 text tokens, quadruple the previous 2048-token limit. It can handle up to six PNG or JPEG images per request, videos up to 120 seconds long, and PDF documents up to six pages. A notable feature is Gemini Embedding2's native support for audio processing, eliminating the need for speech-to-text conversion and avoiding potential information loss from transcription. Google also introduced "interleaved input" technology, enabling developers to combine multiple modalities in one request—like mixing images with descriptive text—to better capture the semantic relationships between them.

Architecturally, the model continues to employ Matryoshka Representation Learning (MRL). This technique uses a hierarchical structure to dynamically adjust vector dimensions. The default embedding dimension is 3072, with optional configurations of 1536 and 768 available, giving developers flexibility to balance retrieval accuracy with storage efficiency.
Google's benchmark results indicate that Gemini Embedding2 delivers leading performance across text, image, video, and speech tasks. For instance, in text-video retrieval, it scores 68.8, outperforming Amazon Nova2Multimodal Embeddings (60.3) and Voyage Multimodal3.5 (55.2). In text-image comparison, it achieves a score of 93.4, significantly ahead of Amazon's model score of 84.0.
Gemini Embedding2 is currently accessible to developers via the Gemini API and Vertex AI. It integrates with popular frameworks and vector databases like LangChain, LlamaIndex, Haystack, Weaviate, Qdrant, ChromaDB, and Vector Search. To help developers get started, Google provides interactive Colab notebooks and lightweight multimodal semantic search demonstrations.

The competition in multimodal embedding is heating up. Notably, in late February of this year, the AI search engine Perplexity released its open-source embedding models, pplx-embed-v1 and pplx-embed-context-v1.
MiniMax Unveils 10x Team Program to Incentivize Global AI Experts
MiniMax (Xiyu Technology), the General Artificial Intelligence Lab, has officially launched "10x Team," a global talent collaboration initiative. This program aims to recruit top experts across industries to explore the deep application of large mode
South Korea Breaks Ground on National AI Computing Center, Investing 2.5 Trillion Won with 2028 Target
South Korean outlet EtNews reports that groundbreaking for the Korea AI Computing Center (KOACC) took place on August 3 at the Solar City data center park in Sunan, Jeollanam-do. Backed by a total investment of 2.5 trillion KRW (roughly 11.838 billio
Six Tech Giants Back Linux Foundation With $12.5M to Tackle AI Vulnerability Noise
To tackle the flood of low-quality security reports produced by AI automation tools, six major tech companies—Anthropic, Amazon (AWS), GitHub, Google, Microsoft, and OpenAI—have collectively contributed $12.5 million in funding to Linux Foundation in





Home






