Wikidata serves as a sister project to Wikipedia, collecting images, texts, keywords, and other information about various topics. For example, for the late English writer Douglas Adams, known for his 1979 work The Hitchhiker's Guide to the Galaxy, you will find not only basic information from his Wikipedia page, but also details such as his zodiac sign, Pisces, or the number 13230702, under which his books are cataloged in libraries around the world. This information is available on the Wikidata website under the identifier Q42, and for machines in formats such as JSON (a data format).
Wikidata contains 19 million entries that volunteers carefully compile. This data is stored not only for people, but also for machines, enabling fast searches and use in various applications.
A new AI-friendly database
Wikimedia Deutschland, the German branch of the Wikimedia Foundation that oversees Wikidata, has launched the Wikipedia Embedding Project. This Berlin-based team spent an entire year using a large language model to convert structured data from Wikidata into vectors. These vectors capture the context and meaning surrounding each entry. For example, Douglas Adams would be associated with the entry "human" and with the titles of his books, as explained by Lydia Pintscher, Wikidata portfolio lead.
This vector form resembles a graph with points and connecting lines, making it easier for large language models to process information. The database was created from data captured through September 18, 2024, and uses a model from Jina AI. The storage infrastructure is provided free of charge by DataStax, which is owned by IBM.
Philippe Saadé, AI project manager at Wikidata, explained that vectors allow artificial intelligence systems to retrieve not only the information itself, but also the context surrounding it. The team is now waiting for feedback from developers before adding data from the past year. According to him, minor changes to existing entries will not affect the database's overall usefulness because the vectors capture the general concept of an entry.
Benefits for developers
The project's goal is to level the playing field for developers outside major technology companies. Large companies such as OpenAI or Anthropic can vectorize Wikidata themselves, but smaller teams benefit the most. Lydia Pintscher emphasized that this gives smaller projects a chance to succeed. For example, the Govdirectory platform uses Wikidata to find the social media accounts and email addresses of public officials around the world.
Thanks to vectors and the new database, the data can be more easily integrated into chatbots or other artificial intelligence applications. This helps systems such as ChatGPT better process lesser-known topics that are not as widely covered online. Instead of waiting for models to be retrained, data can be added directly, improving accuracy.
The user interface will remain the same – Wikipedia will not become a chatbot, the project leads assure. The change primarily concerns the backend, where developers now have easier access to the data.
Supporting knowledge
Wikidata offers stable identifiers such as QIDs for each entry, ensuring unambiguous searches. The data is verified by volunteers and linked to external sources, reducing duplication and improving integration. A new RESTful API (programming interface) speeds up real-time access to data, making it ideal for frameworks such as LangChain.
These improvements make Wikidata a trusted source for artificial intelligence, reducing errors such as model hallucinations. The project supports open-source development and helps smaller teams build tools based on high-quality data.
The team plans further updates based on feedback. The database now contains vectors from 19 million entries, opening the door to applications ranging from search engines to virtual assistants. Lydia Pintscher sees this as a way to ensure that artificial intelligence reflects real knowledge, not just popular trends on the web.
Source: theverge.com



