Gemini 3 Pro Is Currently the Most Reliable AI, but Hallucinations Remain a Problem

Gemini 3 Pro Is Currently the Most Reliable AI, but Hallucinations Remain a Problem

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
4. 12. 2025
4 minutes reading
Gemini 3 Pro Is Currently the Most Reliable AI, but Hallucinations Remain a Problem

New tests are emerging in the field of artificial intelligence to evaluate how well models handle facts. A new benchmark from Artificial Analysis tested 40 large language models, and the results are not very encouraging. Only four of them achieved a positive score in an index called the Omniscience Index, which ranges from -100 to +100. Google’s Gemini 3 Pro model ranked first with 13 points. It was followed by Claude 4.1 Opus with 4.8 points, then GPT-5.1 and Grok 4. This index measures how reliably models provide correct information across different domains.

Omniscience Index results
Omniscience Index results

Gemini 3 Pro achieved such a strong lead mainly due to its accuracy, which stands at 53%. That is 14 points higher than the previous leader, Grok 4. Researchers at Artificial Analysis explain this by noting that accuracy is related to model size—larger models such as Gemini 3 Pro have better factual coverage. A score of 0 would mean that the model provides correct and incorrect answers equally often. The benchmark includes 6,000 questions across 42 topics in six domains: business, humanities and social sciences, health, law, software engineering, and science and mathematics. The questions come from trusted academic and industry sources and were generated automatically by an AI agent.

Hallucinations are AI’s main problem

Although Gemini 3 Pro leads in accuracy, its greatest weakness remains hallucinations—meaning that the model often provides incorrect answers with great confidence instead of admitting that it does not know something. Its hallucination rate reaches 88%, the same as the Gemini 2.5 Pro and Gemini 2.5 Flash models. By comparison, GPT-5.1 (high) has a rate of 81% and Grok 4 has 64%. This means that among all incorrect answers, a large proportion involve the model inventing facts rather than admitting that it does not know.

Claude 4.1 Opus achieved an accuracy of 36%, but with one of the lowest hallucination rates, which had previously secured it first place. The researchers emphasize that high hallucination rates are a problem across all tested models. Here, hallucination refers to the proportion of incorrect answers among all unsuccessful attempts, highlighting the models’ excessive confidence.

Hallucination results in the Omniscience Index
Hallucination results in the Omniscience Index

Wrong answer? You’ll be penalized!

This benchmark differs from conventional tests in that it penalizes incorrect answers just as strongly as it rewards correct ones. Conventional methods often encourage guessing, which increases hallucinations. In the Omniscience Index, models receive no points for admitting that they do not know, but they are not penalized either—in contrast, incorrect answers result in substantial deductions. This is intended to encourage models to be cautious.

The results divide the models into four groups: those with extensive knowledge and high reliability, such as Claude 4.1 Opus; those with knowledge but low reliability, such as Claude 4.5 Haiku; those with limited knowledge but consistent reliability, such as GPT-5.1; and small models lacking both knowledge and reliability, such as OpenAI’s lightweight gpt-oss. A detailed breakdown by domain is not available for Gemini 3 Pro.

Older Llama model surprises with strong results

General intelligence does not necessarily imply factual reliability. Models such as Minimax M2 or gpt-oss-120b (high) perform well in Artificial Analysis’s broader intelligence index, but fail in the Omniscience Index due to high hallucination rates. Conversely, the older Llama-3.1-405B model achieved a good score, even though it usually lags behind newer models in overall evaluations.

No model consistently excelled across all six domains. Claude 4.1 Opus led in law, software engineering, and the humanities, GPT-5.1.1 in business questions, and Grok 4 in health and science. These differences mean that overall ratings may conceal important gaps.

Results across six domains in the Omniscience Index
Results across six domains in the Omniscience Index

Model size does not always mean better reliability

Larger models generally achieve higher accuracy, but not necessarily lower hallucination rates. Several smaller models, such as Nvidia’s Nemotron Nano 9B V2 or Llama Nemotron Super 49B v1.5, outperformed much larger competitors in the Omniscience Index. Artificial Analysis confirmed that accuracy is related to size, but hallucinations are not. Therefore, despite its high accuracy, Gemini 3 Pro still frequently hallucinates.

In terms of cost, Claude 4.5 Haiku stands out with a better score than more expensive models such as GPT-5.1 (high) or Kimi K2 Thinking. The researchers published 10% of the questions as a public dataset on Hugging Face, while the remainder remains private to prevent contamination of training data. A related study revealed flaws in existing benchmarks, such as unclear definitions of terms or insufficient statistical validation.

Source: the-decoder.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok