Home
Microsoft Bing Team Open Sources 27B Embedding Model Harrier, Leading Multilingual Benchmarks
On April 7, Microsoft's Bing team announced the open-source release of a new word embedding model series called "Harrier," designed to reshape the underlying logic of global search, retrieval, and AI agents. The Harrier series includes three versions, with the flagship 27B model surpassing major proprietary models from OpenAI, Amazon, and Google Gemini in the multilingual MTEB v2 benchmark, securing the top spot.

The technical foundation of this model reflects strong industrial standards: Harrier supports over 100 languages with a context window of up to 32,000 tokens. For training, Microsoft used over 2 billion real-world examples and incorporated synthetic data from GPT-5 for reinforcement learning. This high-quality data mix gives Harrier a distinct advantage in understanding complex contexts and handling long texts. Alongside the full 27B-parameter version, Microsoft also released smaller 0.6B and 2.7B variants to accommodate different computing environments, all open-sourced on Hugging Face under the MIT license.
Embedding models are essential for organizing and retrieving information in AI systems, and their performance directly impacts the accuracy of RAG (Retrieval-Augmented Generation) systems. Microsoft intends to integrate this technology into the Bing search engine and new AI agent grounding services. As AI moves toward more autonomous, multi-step tasks, the open-sourcing of Harrier not only gives developers a high-performance alternative to proprietary models but also marks a significant milestone for the open-source ecosystem in semantic representation, further accelerating AI agent deployment in multilingual global environments.
Related article
ByteDance’s Seed launches global campus drive, offering virtual shares to win top large model talent
In the competitive landscape of large language models, securing top-tier talent remains the most critical strategic asset.On April 1st, ByteDance announced the launch of its Seed global campus recruitment initiative, part of its large model talent de
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
Related Special Topic Recommendations
Comments (1)
0/500
Finally, an open-source embedding model that actually challenges the big players! Harrier's 27B parameters look promising for multilingual tasks, especially if it improves retrieval accuracy without the usual proprietary lock-in. Microsoft is pushing boundaries here, but let's see how it performs in real-world edge cases. Excited to test it out on some niche datasets! 🚀
On April 7, Microsoft's Bing team announced the open-source release of a new word embedding model series called "Harrier," designed to reshape the underlying logic of global search, retrieval, and AI agents. The Harrier series includes three versions, with the flagship 27B model surpassing major proprietary models from OpenAI, Amazon, and Google Gemini in the multilingual MTEB v2 benchmark, securing the top spot.

The technical foundation of this model reflects strong industrial standards: Harrier supports over 100 languages with a context window of up to 32,000 tokens. For training, Microsoft used over 2 billion real-world examples and incorporated synthetic data from GPT-5 for reinforcement learning. This high-quality data mix gives Harrier a distinct advantage in understanding complex contexts and handling long texts. Alongside the full 27B-parameter version, Microsoft also released smaller 0.6B and 2.7B variants to accommodate different computing environments, all open-sourced on Hugging Face under the MIT license.
Embedding models are essential for organizing and retrieving information in AI systems, and their performance directly impacts the accuracy of RAG (Retrieval-Augmented Generation) systems. Microsoft intends to integrate this technology into the Bing search engine and new AI agent grounding services. As AI moves toward more autonomous, multi-step tasks, the open-sourcing of Harrier not only gives developers a high-performance alternative to proprietary models but also marks a significant milestone for the open-source ecosystem in semantic representation, further accelerating AI agent deployment in multilingual global environments.
ByteDance’s Seed launches global campus drive, offering virtual shares to win top large model talent
In the competitive landscape of large language models, securing top-tier talent remains the most critical strategic asset.On April 1st, ByteDance announced the launch of its Seed global campus recruitment initiative, part of its large model talent de
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
Finally, an open-source embedding model that actually challenges the big players! Harrier's 27B parameters look promising for multilingual tasks, especially if it improves retrieval accuracy without the usual proprietary lock-in. Microsoft is pushing boundaries here, but let's see how it performs in real-world edge cases. Excited to test it out on some niche datasets! 🚀











