Research led by scientists from Aalto University reveals an interesting phenomenon. When people use large language models such as ChatGPT to solve complex tasks, their actual performance improves. For example, in logical reasoning tests from the Law School Admission Test (LSAT), participants scored an average of three points higher than the general population. This means that AI genuinely helps improve results in demanding cognitive tasks. The study, conducted by Daniela Fernandes, Steeven Villa, Salla Nicholls, Otso Haavisto, Daniel Buschek, Albrecht Schmidt, Thomas Kosch, Chenxinran Shen, and Robin Welsch, involved two large groups of people. The first part included 246 participants, while the second included 452.
Participants were asked to solve 20 problems from the LSAT, a test used for admission to law schools in the US. Those who used AI achieved better results, but they also significantly overestimated how well they had performed. On average, they thought they had scored four points higher than they actually had. This suggests that although AI increases effectiveness, people lose the ability to accurately assess their own contribution.
The problem of overestimation and the Dunning-Kruger effect
One of the key findings is that the traditional Dunning-Kruger effect, in which lower-performing individuals overestimate their abilities while higher-performing individuals underestimate theirs, completely disappears when AI is used. Under normal conditions without AI, this effect holds true—people with lower performance believe they are better than they actually are. But when AI enters the picture, this bias levels out. All participants, regardless of their ability level, overestimated their performance to the same extent.
The researchers used a computational model to analyze individual differences. It showed that AI levels out both cognitive performance and self-assessment ability. This means that AI helps lower-performing individuals achieve better results, but at the same time leads everyone to have an overly optimistic view of their abilities. In the second study, where accurate self-assessment was incentivized with a financial reward, the same pattern emerged. People still overestimated their performance even though they had a reason to reflect more deeply.
The role of AI literacy in self-assessment
An interesting paradox emerged among people with greater knowledge of AI. The researchers measured AI literacy using the SNAIL scale (Scale for the Assessment of Non-Experts' AI Literacy), developed by M.C. Laupichler and colleagues. Those with better technical knowledge of AI were more confident, but their performance estimates were less accurate. Higher confidence correlated with lower self-assessment accuracy. This means that people who understand AI tend to overestimate their success even more than those with lower AI literacy.
In practice, most participants interacted with ChatGPT only minimally—often submitting just one prompt per question. They copied the problem into the AI, accepted its answer without further verification, and trusted the system blindly. This approach, known as cognitive offloading, means leaving all processing to AI, which limits their ability to reflect on their own mistakes.
Implications for the everyday use of AI
The research highlights the risks associated with placing too much trust in AI. People often use AI for complex tasks such as logical reasoning but fail to realize that their actual contribution is smaller than they think. This may lead to excessive reliance on such systems, reducing their own cognitive skills in the long term. For example, in the first study, performance improved, but metacognitive accuracy declined—people were less able to distinguish between correct and incorrect answers than they would have been without AI.
The researchers compared their results with data from a previous 2021 study by Jansen et al., in which participants solved the same tasks without AI. The Dunning-Kruger effect was present there, but the situation changed with AI. This suggests that AI may level out individual differences, but at the cost of accurate self-assessment.
Ways to improve interaction with AI
To minimize these problems, the authors suggest designing better interfaces for AI systems. For example, AI could ask users to explain their reasoning in greater detail, which would encourage critical thinking. The second study, which required more interactions with AI, showed that deeper engagement can help, but it still did not fully improve metacognition.
The study therefore concludes that although AI increases productivity, it is important to develop tools that encourage reflection. Without them, we risk becoming overly dependent on technology without being aware of our limitations. The research was supported by the Finnish Doctoral Program in Artificial Intelligence and the European Research Council's AmplifAI project.



