Wikipedia partners with AI firms to enhance data access

Wikimedia Deutschland unveiled a new database this Wednesday designed to enhance AI models' access to Wikipedia's extensive knowledge repository.
Dubbed the Wikidata Embedding Project, this initiative leverages vector-based semantic search technology—enabling computers to grasp word meanings and relationships—applied to Wikipedia’s vast network of nearly 120 million entries across its sister platforms.
With added compatibility for the Model Context Protocol (MCP), a framework facilitating AI-data source communication, the project improves how large language models interact with and retrieve information through natural language queries.
Wikimedia’s German division spearheaded this effort alongside Jina.AI, a neural search specialist, and IBM-owned DataStax, which specializes in real-time training data.
While Wikidata has long provided machine-readable data from Wikimedia properties, previous tools were limited to keyword searches and SPARQL queries. This upgraded system enhances retrieval-augmented generation (RAG) capabilities, enabling AI developers to anchor their models in Wikipedia’s editor-verified knowledge base.
The database structures data with rich semantic context. For example, searching "scientist" returns lists of notable nuclear scientists and Bell Labs researchers, along with multilingual translations, curated Wikimedia images, and related terms like "researcher" and "scholar."
Accessible via Toolforge, Wikidata will host a developer webinar on October 9th to showcase the platform’s potential.
Connect with 10,000+ tech and VC pioneers at Disrupt 2025
Join industry giants like Netflix, Box, a16z, ElevenLabs, and Vinod Khosla across 200+ sessions packed with startup growth strategies and tech insights. Secure your early-bird ticket today—save up to $444 before general admission opens.
Connect with 10,000+ tech and VC pioneers at Disrupt 2025
Join industry giants like Netflix, Box, a16z, ElevenLabs, and Vinod Khosla across 200+ sessions packed with startup growth strategies and tech insights. Secure your early-bird ticket today—save up to $444 before general admission opens.
This launch arrives as AI developers increasingly seek premium data sources for model refinement. Modern training systems—now intricate ecosystems rather than simple datasets—still demand meticulously curated information, especially for accuracy-critical applications. Wikipedia’s fact-checked content offers a stark advantage over bulk datasets like Common Crawl.
The quest for quality data carries risks: Anthropic recently proposed a $1.5 billion settlement after authors sued over unauthorized use of their works for training.
Wikidata’s AI project lead Philippe Saadé highlighted its independence: "This proves powerful AI can thrive beyond corporate silos—open, collaborative, and built for public benefit," he told press.
Related article
Ollie bets privacy focus to win AI assistant race
To be genuinely helpful, an AI assistant must understand its user deeply. Ollie, a personal assistant designed for daily life, operates on the premise that this doesn’t require surrendering your data or compromising your privacy.While certain enterpr
How AI LIVE: London Will Explore AI & Industrial Automation
The summit will convene C-suite executives from around the globe to address pressing challenges in global industries, ranging from AI-driven disruption to economic volatility.AI LIVE: The London Summit will gather over 2,000 international leaders und
Anthropic Enters AI Legal Tech Market as Competition Intensifies
Anthropic unveiled a suite of new chatbot capabilities on Tuesday, aimed at delivering automated support to legal practices. These enhancements expand upon Claude for Legal, the firm-specific platform introduced earlier this year, by adding specializ
Related Special Topic Recommendations
Comments (3)
0/500
So they're basically selling training data to AI companies? As long as the knowledge stays accessible, I'm cool with it. 🤔
Das ist ein wirklich cleverer Schachzug von Wikipedia! Vektorsuche in ihren riesigen Datenbeständen könnte die Qualität von KI-Ausgaben enorm verbessern und vielleicht endlich mit den Halluzinationen aufräumen. Hoffentlich bleibt der Zugang aber transparent und für alle fair, damit nicht nur die großen Tech-Konzerne profitieren. Die deutsche Wikimedia-Abteilung zeigt mal wieder, dass sie vorne mitmischt. 💡

Wikimedia Deutschland unveiled a new database this Wednesday designed to enhance AI models' access to Wikipedia's extensive knowledge repository.
Dubbed the Wikidata Embedding Project, this initiative leverages vector-based semantic search technology—enabling computers to grasp word meanings and relationships—applied to Wikipedia’s vast network of nearly 120 million entries across its sister platforms.
With added compatibility for the Model Context Protocol (MCP), a framework facilitating AI-data source communication, the project improves how large language models interact with and retrieve information through natural language queries.
Wikimedia’s German division spearheaded this effort alongside Jina.AI, a neural search specialist, and IBM-owned DataStax, which specializes in real-time training data.
While Wikidata has long provided machine-readable data from Wikimedia properties, previous tools were limited to keyword searches and SPARQL queries. This upgraded system enhances retrieval-augmented generation (RAG) capabilities, enabling AI developers to anchor their models in Wikipedia’s editor-verified knowledge base.
The database structures data with rich semantic context. For example, searching "scientist" returns lists of notable nuclear scientists and Bell Labs researchers, along with multilingual translations, curated Wikimedia images, and related terms like "researcher" and "scholar."
Accessible via Toolforge, Wikidata will host a developer webinar on October 9th to showcase the platform’s potential.
Connect with 10,000+ tech and VC pioneers at Disrupt 2025
Join industry giants like Netflix, Box, a16z, ElevenLabs, and Vinod Khosla across 200+ sessions packed with startup growth strategies and tech insights. Secure your early-bird ticket today—save up to $444 before general admission opens.
Connect with 10,000+ tech and VC pioneers at Disrupt 2025
Join industry giants like Netflix, Box, a16z, ElevenLabs, and Vinod Khosla across 200+ sessions packed with startup growth strategies and tech insights. Secure your early-bird ticket today—save up to $444 before general admission opens.
This launch arrives as AI developers increasingly seek premium data sources for model refinement. Modern training systems—now intricate ecosystems rather than simple datasets—still demand meticulously curated information, especially for accuracy-critical applications. Wikipedia’s fact-checked content offers a stark advantage over bulk datasets like Common Crawl.
The quest for quality data carries risks: Anthropic recently proposed a $1.5 billion settlement after authors sued over unauthorized use of their works for training.
Wikidata’s AI project lead Philippe Saadé highlighted its independence: "This proves powerful AI can thrive beyond corporate silos—open, collaborative, and built for public benefit," he told press.
Ollie bets privacy focus to win AI assistant race
To be genuinely helpful, an AI assistant must understand its user deeply. Ollie, a personal assistant designed for daily life, operates on the premise that this doesn’t require surrendering your data or compromising your privacy.While certain enterpr
How AI LIVE: London Will Explore AI & Industrial Automation
The summit will convene C-suite executives from around the globe to address pressing challenges in global industries, ranging from AI-driven disruption to economic volatility.AI LIVE: The London Summit will gather over 2,000 international leaders und
Anthropic Enters AI Legal Tech Market as Competition Intensifies
Anthropic unveiled a suite of new chatbot capabilities on Tuesday, aimed at delivering automated support to legal practices. These enhancements expand upon Claude for Legal, the firm-specific platform introduced earlier this year, by adding specializ
So they're basically selling training data to AI companies? As long as the knowledge stays accessible, I'm cool with it. 🤔
Das ist ein wirklich cleverer Schachzug von Wikipedia! Vektorsuche in ihren riesigen Datenbeständen könnte die Qualität von KI-Ausgaben enorm verbessern und vielleicht endlich mit den Halluzinationen aufräumen. Hoffentlich bleibt der Zugang aber transparent und für alle fair, damit nicht nur die großen Tech-Konzerne profitieren. Die deutsche Wikimedia-Abteilung zeigt mal wieder, dass sie vorne mitmischt. 💡





Home






