New tests are emerging in the field of artificial intelligence to evaluate how well models handle facts. A new benchmark from Artificial Analysis tested 40 large language models, and the results are not very encouraging. Only four of them achieved a positive score in an index called the Omniscience Index, which ranges from -100 to +100. Google’s Gemini 3 Pro model ranked first with 13 points. It was followed by Claude 4.1 Opus with 4.8 points, then GPT-5.1 and Grok 4. This index measures how reliably models provide correct information across different domains.
Gemini 3 Pro achieved such a strong lead mainly due to its accuracy, which stands at 53%. That is 14 points higher than the previous leader, Grok 4. Researchers at Artificial Analysis explain this by noting that accuracy is related to model size—larger models such as Gemini 3 Pro have better factual coverage. A score of 0 would mean that the model provides correct and incorrect answers equally often. The benchmark includes 6,000 questions across 42 topics in six domains: business, humanities and social sciences, health, law, software engineering, and science and mathematics. The questions come from trusted academic and industry sources and were generated automatically by an AI agent.
Hallucinations are AI’s main problem
Although Gemini 3 Pro leads in accuracy, its greatest weakness remains hallucinations—meaning that the model often provides incorrect answers with great confidence instead of admitting that it does not know something. Its hallucination rate reaches 88%, the same as the Gemini 2.5 Pro and Gemini 2.5 Flash models. By comparison, GPT-5.1 (high) has a rate of 81% and Grok 4 has 64%. This means that among all incorrect answers, a large proportion involve the model inventing facts rather than admitting that it does not know.
Claude 4.1 Opus achieved an accuracy of 36%, but with one of the lowest hallucination rates, which had previously secured it first place. The researchers emphasize that high hallucination rates are a problem across all tested models. Here, hallucination refers to the proportion of incorrect answers among all unsuccessful attempts, highlighting the models’ excessive confidence.
Wrong answer? You’ll be penalized!
This benchmark differs from conventional tests in that it penalizes incorrect answers just as strongly as it rewards correct ones. Conventional methods often encourage guessing, which increases hallucinations. In the Omniscience Index, models receive no points for admitting that they do not know, but they are not penalized either—in contrast, incorrect answers result in substantial deductions. This is intended to encourage models to be cautious.
The results divide the models into four groups: those with extensive knowledge and high reliability, such as Claude 4.1 Opus; those with knowledge but low reliability, such as Claude 4.5 Haiku; those with limited knowledge but consistent reliability, such as GPT-5.1; and small models lacking both knowledge and reliability, such as OpenAI’s lightweight gpt-oss. A detailed breakdown by domain is not available for Gemini 3 Pro.
Older Llama model surprises with strong results
General intelligence does not necessarily imply factual reliability. Models such as Minimax M2 or gpt-oss-120b (high) perform well in Artificial Analysis’s broader intelligence index, but fail in the Omniscience Index due to high hallucination rates. Conversely, the older Llama-3.1-405B model achieved a good score, even though it usually lags behind newer models in overall evaluations.
No model consistently excelled across all six domains. Claude 4.1 Opus led in law, software engineering, and the humanities, GPT-5.1.1 in business questions, and Grok 4 in health and science. These differences mean that overall ratings may conceal important gaps.
Model size does not always mean better reliability
Larger models generally achieve higher accuracy, but not necessarily lower hallucination rates. Several smaller models, such as Nvidia’s Nemotron Nano 9B V2 or Llama Nemotron Super 49B v1.5, outperformed much larger competitors in the Omniscience Index. Artificial Analysis confirmed that accuracy is related to size, but hallucinations are not. Therefore, despite its high accuracy, Gemini 3 Pro still frequently hallucinates.
In terms of cost, Claude 4.5 Haiku stands out with a better score than more expensive models such as GPT-5.1 (high) or Kimi K2 Thinking. The researchers published 10% of the questions as a public dataset on Hugging Face, while the remainder remains private to prevent contamination of training data. A related study revealed flaws in existing benchmarks, such as unclear definitions of terms or insufficient statistical validation.
Source: the-decoder.com



