AI Models Break Records on CFA Exams

AI Models Break Records on CFA Exams

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
19. 12. 2025
4 minutes reading
AI Models Break Records on CFA Exams

What Exactly Is the CFA?

The CFA, or Chartered Financial Analyst, is a prestigious certification for investment and finance professionals. The program is divided into three levels, each testing different skills. Level I focuses on foundational knowledge through standalone multiple-choice questions. Level II examines application and analysis using sets of questions based on case studies. Level III then assesses complex synthesis and portfolio construction through a combination of multiple-choice questions and questions requiring written responses. It is a demanding process that requires accurate calculations, high-quality analysis, and sound reasoning.

The research examines how modern artificial intelligence models perform on these exams. The researchers used a set of mock exams containing a total of 980 questions: three exams for Level I (540 questions), two for Level II (176 questions), and three for Level III (264 questions). These exams came from official CFA Institute materials and sources such as AnalystPrep, reflecting the 2024 and 2025 curriculum updates, including the new specialized pathways for Level III.

How Was the Testing Conducted?

The researchers first reproduced the results of older models, such as ChatGPT (GPT-3.5-turbo), GPT-4, and GPT-4o, to establish a comparative baseline. They then tested advanced models such as GPT-5, Gemini 3.0 Pro, Gemini 2.5 Pro, Grok 4, Claude Opus 4.1, and DeepSeek-V3.1. The tests were conducted in two modes: zero-shot, in which the model received the question directly without additional instructions, and chain-of-thought, in which it was prompted to reason step by step and provide an explanation.

They used precise passing criteria for the assessment. At Level I, the model had to achieve at least 60% in each topic area and 70% overall. At Level II, the thresholds were 50% in each area and 60% overall. At Level III, an average score of 63% across the combination of multiple-choice questions and written responses was sufficient. The written responses were evaluated by an automated system based on models such as o4-mini, using a rubric from AnalystPrep.

The tests covered ten main topics for Levels I and II, including quantitative methods, economics, financial reporting, corporate issuers, equity investments, fixed income, derivatives, alternative investments, and portfolio management. For Level III, they focused on areas such as asset allocation, portfolio construction, performance measurement, derivatives and risk management, ethical standards, and specialized pathways.

Results: AI Exceeded Expectations

Older models such as ChatGPT achieved scores ranging from approximately 58.9% to 68.4% at Level I, which was not enough to pass. GPT-4 passed Level I with scores ranging from 73.3% to 80.9%, but failed Level II. GPT-4o passed Levels I and II with 90.6% and 73.9%, respectively, and even passed Level III with 66.7% on the written questions.

However, the advanced models dominated. Gemini 3.0 Pro achieved a record 97.6% at Level I in zero-shot mode. GPT-5 led at Level II with 94.3%. At Level III, Gemini 2.5 Pro excelled on multiple-choice questions with 86.4%, while Gemini 3.0 Pro achieved 92.0% on written responses. All of these models passed every level, with the overall performance ranking as follows: Gemini 3.0 Pro, Gemini 2.5 Pro, GPT-5, Grok 4, Claude Opus 4.1, and DeepSeek-V3.1.

Model success rates
Model success rates.

Chain-of-thought mode improved the performance of older models, but its effect on newer models was mixed—for example, Gemini 3.0 Pro saw a slight decline on multiple-choice questions but a major improvement on written responses. Errors were concentrated mainly in ethical standards, where models such as GPT-5 had error rates of 17–21% at Level II.

Details About the Models and Data Tested

The models were tested with the temperature set to 0 to minimize randomness, and the results included averages with deviations. For example, GPT-5 used the identifier gpt-5-preview dated August 7, 2025, Gemini 3.0 Pro used gemini-3-pro-preview dated November 18, 2025, and Grok 4 used grok-4 dated July 9, 2025.

The data was carefully selected to reflect the current curriculum. For example, at Level III, it included specialized questions on areas such as portfolio management, private markets, and private wealth. A comparison with a previous study confirmed that the new exams have a similar topic distribution, but a lower proportion of computational questions due to curriculum updates.

The research acknowledges limitations, such as the risk of data contamination from the models' training datasets or potential bias in the automated evaluation of written responses, where longer texts may be favored. For Level III, the researchers used third-party providers such as AnalystPrep, which may not fully correspond to the official exams.

These findings suggest that artificial intelligence has reached a level at which it can handle knowledge and synthesis comparable to that of experienced financial analysts.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok