How Far Can We Scale Reasoning Models?

How Far Can We Scale Reasoning Models?

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
21. 5. 2025
4 minutes reading
How Far Can We Scale Reasoning Models?

How Far Can Reasoning-Focused Artificial Intelligence Models Scale? An Analysis of Future Limits and Possibilities

The Rapid Rise of Reasoning Models Will Soon Reach Its Limits

Artificial intelligence models specialized in reasoning, such as OpenAI o3, have experienced a dramatic increase in computing power in recent months. The current trajectory shows that the computing capacity of these models is increasing approximately tenfold every few months, leading to significant improvements in areas such as mathematics, science, and programming. However, this exponential growth rate is running up against physical limits - if the current trend continued, training reasoning-focused models would reach the limit of the total available computing power within approximately one year. Once this milestone is reached, further scaling will be limited by the overall growth rate of hardware infrastructure, which is currently approximately fourfold per year, significantly slower than the current explosive growth.

Scaling Laws and Performance Curves

While rigorous scaling laws specific to training reasoning-focused models are still taking shape (unlike the well-established laws for pretraining), early evidence suggests that performance on tasks such as mathematics and programming grows log-linearly with an increasing number of steps in specialized reasoning-focused training. This pattern mirrors classic neural scaling trends. This means we can expect continued dramatic improvements as additional computing resources are invested in specialized reasoning training phases through reinforcement learning (RL).

Limiting Factors and Bottlenecks

The main limiting factor for advanced reasoning models is the total available computing power. Once reasoning model training reaches the scale of a model's total pretraining runs (which already cost tens of millions of dollars per run), further acceleration will depend on broader advances in artificial intelligence infrastructure or significant new investments. Some laboratories may begin to encounter diminishing returns on investment if they fall behind algorithmic frontiers - improvements resulting solely from scaling may reach their limit unless accompanied by better methods or architectures. The serial nature of reinforcement learning (RL) is also a significant limitation. The reinforcement learning phases used for advanced reasoning are more serial than standard pretraining, making them more difficult to parallelize across massive GPU clusters - this represents a practical limit on how quickly RL-based reasoning can scale.

Synthetic Data and Compute at Test Time

Research laboratories are increasingly using synthetic data generated by powerful reasoning models to train both future reasoning models and models not focused on reasoning. This approach enables more efficient use of computing power without simply increasing the number of parameters or the size of datasets. It could help extend scalable progress even when traditional methods reach their limits. This strategy will likely become even more important in the coming years as available computing power becomes the primary limiting factor.

Short-Term and Long-Term Outlook

In the short term (less than one year), we can expect rapid improvements driven by increased investment in reinforcement learning-based post-training. The growth rate of computing power should remain as high as tenfold every few months, with the possibility of dramatic jumps in performance. The main limiting factor will remain the training budget and available computing power, while synthetic data will play an increasingly important role. In the medium and long term (more than one year), once total available computing power becomes the limiting factor, progress will slow but continue at a steadier pace tied to hardware advances. The growth rate of computing power will slow toward an approximately fourfold annual increase. Performance improvements will be steady but slower, with hardware and infrastructure limits serving as the main constraints. Synthetic data will become absolutely essential for efficient scaling.

The Future of Reasoning Models

Reasoning models may continue their rapid rise for approximately another year before encountering hard resource constraints. After that, their growth rate will align with broader trends in artificial intelligence hardware, unless significant breakthroughs in efficiency or investment occur. Although jumps of several orders of magnitude within current paradigms seem unlikely over this horizon, significant progress will remain achievable through smarter algorithms and synthetic data generation - even after gains from simply increasing computing power begin to decline. Although there is still considerable room before reaching fundamental physical or economic barriers that would halt meaningful gains, order-of-magnitude differences compared with today's largest models seem unlikely without transformative breakthroughs or massive new investments. However, the pace of innovation in this field remains highly unpredictable, which is why the future development of reasoning models remains one of the most fascinating questions in artificial intelligence research.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok