Why AI Models Fail at Writing High-Performance Code Despite Growing Popularity

Why AI Models Fail at Writing High-Performance Code Despite Growing Popularity

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
21. 5. 2025
3 minutes reading
Why AI Models Fail at Writing High-Performance Code Despite Growing Popularity

Why AI Models Fail at Writing High-Performance Code Despite Growing Popularity

Large language models (LLMs) have become a significant part of the development process at many companies, yet according to recent surveys, they still fail to create code that is truly high-performance and optimized. This reality raises important questions about the future of programming, as more and more companies integrate AI tools into their development cycles.

Statistics from the Infobip Shift Conference in Miami

Saurabh Misra, CEO of CodeFlash, recently presented survey results at the Infobip Shift conference that shed light on the true capabilities of LLMs in code generation. According to his findings, while technology giants such as Microsoft and Google already rely on artificial intelligence to create a significant share of their codebase (Microsoft 25% and Google as much as 30%), there is a fundamental difference between functional code and code that is truly high-performance and optimized. The situation is even more pronounced at some startups, where an incredible 95% of all new code comes from AI tools. Misra emphasizes, however, that this high adoption rate does not automatically mean better software performance. In fact, the opposite is often true – language models can quickly write code that works, but they often fail to optimize it for speed, memory usage, or other critical performance metrics.

Shortcomings in Code Writing

The survey highlights a fundamental difference between LLMs’ ability to generate syntactically correct code and their ability to create code that is actually efficient. While current models are relatively reliable at producing code that compiles and performs basic functions, they lack the more sophisticated understanding of optimization that experienced human developers naturally apply. Human programmers constantly consider various edge cases and optimization strategies when writing code, which is something current AI models cannot consistently replicate. One of the main challenges is also assessing the actual performance of generated code. Standard evaluation metrics for LLM outputs usually focus on correctness or logical consistency rather than runtime efficiency or resource usage. Current automated evaluation systems struggle to accurately assess nuanced aspects such as algorithmic complexity or real-world processing speed – a key reason why high-performance programming remains an unattainable goal for most current LLMs.

Humans vs. AI

If we compare the capabilities of human developers and LLMs, we find several fundamental differences. While both often demonstrate a high degree of functional correctness, human developers consistently consider performance aspects, whereas LLMs often overlook them. Human programmers are also context-aware, while LLMs are limited by their training data. Despite the rapid progress in generative AI programming tools and their widespread adoption in the industry, surveys suggest that current LLMs cannot consistently produce high-performance code. They excel at quickly generating working solutions, but typically lack deeper optimization unless they are specifically guided or reviewed by human experts. While generative AI continues to transform the way developers work, the results of these surveys emphasize the ongoing importance of human expertise in creating truly efficient and optimized software. It will therefore be crucial for technology companies to find the right balance between using AI to accelerate development and maintaining high standards of code performance.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok