AI Models Break Records on CFA Exams

AI Models Break Records on CFA Exams

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
19. 12. 2025
4 minutes reading · 6 views
AI Models Break Records on CFA Exams

What Exactly Is the CFA?

The CFA, or Chartered Financial Analyst, is a prestigious certification for investment and finance professionals. The program is divided into three levels, each testing different skills. Level I focuses on foundational knowledge through standalone multiple-choice questions. Level II examines application and analysis using sets of questions based on case studies. Level III then assesses complex synthesis and portfolio construction through a combination of multiple-choice questions and questions requiring written responses. It is a demanding process that requires accurate calculations, high-quality analysis, and sound reasoning.

The research examines how modern artificial intelligence models perform on these exams. The researchers used a set of mock exams containing a total of 980 questions: three exams for Level I (540 questions), two for Level II (176 questions), and three for Level III (264 questions). These exams came from official CFA Institute materials and sources such as AnalystPrep, reflecting the 2024 and 2025 curriculum updates, including the new specialized pathways for Level III.

How Was the Testing Conducted?

The researchers first reproduced the results of older models, such as ChatGPT (GPT-3.5-turbo), GPT-4, and GPT-4o, to establish a comparative baseline. They then tested advanced models such as GPT-5, Gemini 3.0 Pro, Gemini 2.5 Pro, Grok 4, Claude Opus 4.1, and DeepSeek-V3.1. The tests were conducted in two modes: zero-shot, in which the model received the question directly without additional instructions, and chain-of-thought, in which it was prompted to reason step by step and provide an explanation.

They used precise passing criteria for the assessment. At Level I, the model had to achieve at least 60% in each topic area and 70% overall. At Level II, the thresholds were 50% in each area and 60% overall. At Level III, an average score of 63% across the combination of multiple-choice questions and written responses was sufficient. The written responses were evaluated by an automated system based on models such as o4-mini, using a rubric from AnalystPrep.

The tests covered ten main topics for Levels I and II, including quantitative methods, economics, financial reporting, corporate issuers, equity investments, fixed income, derivatives, alternative investments, and portfolio management. For Level III, they focused on areas such as asset allocation, portfolio construction, performance measurement, derivatives and risk management, ethical standards, and specialized pathways.

Results: AI Exceeded Expectations

Older models such as ChatGPT achieved scores ranging from approximately 58.9% to 68.4% at Level I, which was not enough to pass. GPT-4 passed Level I with scores ranging from 73.3% to 80.9%, but failed Level II. GPT-4o passed Levels I and II with 90.6% and 73.9%, respectively, and even passed Level III with 66.7% on the written questions.

However, the advanced models dominated. Gemini 3.0 Pro achieved a record 97.6% at Level I in zero-shot mode. GPT-5 led at Level II with 94.3%. At Level III, Gemini 2.5 Pro excelled on multiple-choice questions with 86.4%, while Gemini 3.0 Pro achieved 92.0% on written responses. All of these models passed every level, with the overall performance ranking as follows: Gemini 3.0 Pro, Gemini 2.5 Pro, GPT-5, Grok 4, Claude Opus 4.1, and DeepSeek-V3.1.

Model success rates
Model success rates.

Chain-of-thought mode improved the performance of older models, but its effect on newer models was mixed—for example, Gemini 3.0 Pro saw a slight decline on multiple-choice questions but a major improvement on written responses. Errors were concentrated mainly in ethical standards, where models such as GPT-5 had error rates of 17–21% at Level II.

Details About the Models and Data Tested

The models were tested with the temperature set to 0 to minimize randomness, and the results included averages with deviations. For example, GPT-5 used the identifier gpt-5-preview dated August 7, 2025, Gemini 3.0 Pro used gemini-3-pro-preview dated November 18, 2025, and Grok 4 used grok-4 dated July 9, 2025.

The data was carefully selected to reflect the current curriculum. For example, at Level III, it included specialized questions on areas such as portfolio management, private markets, and private wealth. A comparison with a previous study confirmed that the new exams have a similar topic distribution, but a lower proportion of computational questions due to curriculum updates.

The research acknowledges limitations, such as the risk of data contamination from the models' training datasets or potential bias in the automated evaluation of written responses, where longer texts may be favored. For Level III, the researchers used third-party providers such as AnalystPrep, which may not fully correspond to the official exams.

These findings suggest that artificial intelligence has reached a level at which it can handle knowledge and synthesis comparable to that of experienced financial analysts.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok