Why Andrej Karpathy Does Not Believe in RL for AI Training—Former OpenAI Employee
Andrej Karpathy, a former researcher at Tesla and OpenAI, shared his long-standing skepticism on X about reinforcement learning (RL) as a core method for training large language models (LLMs). In the post, he describes reward functions in RL as "super sus"—that is, highly unreliable, easy to manipulate, and unsuitable for teaching genuine intellectual problem-solving skills. This position runs counter to that of many major players such as OpenAI, which see RL as a scalable approach for new tasks, even though purely pretrained LLMs have seemingly reached their peak.
An article on The Decoder explores Karpathy's views in detail. According to him, RL works best when there is a clearly right or wrong answer, because the model receives positive feedback for solving problems step by step. This helps LLMs break tasks down into logical steps and makes their reasoning more transparent. However, Karpathy warns that for more complex cognitive tasks, such as intellectual problem-solving, these reward functions are inadequate and can easily be circumvented.
In era of pretraining, what mattered was internet text. You'd primarily want a large, diverse, high quality collection of internet documents to learn from.
— Andrej Karpathy (@karpathy) August 27, 2025
In era of supervised finetuning, it was conversations. Contract workers are hired to create answers for questions, a bit… https://t.co/rR6yYZGgKP
Karpathy Acknowledges the Benefits of RL but Calls for Change
Although Karpathy criticizes RL, he concedes that fine-tuning models using RL is a step forward compared with conventional supervised fine-tuning (SFT), which merely imitates human responses. In his words, RL leads to more sophisticated model behavior, and he expects the method to continue developing significantly. In another post on X, he mentions that RL fine-tuning "will continue to grow substantially."
But according to Karpathy, the real breakthroughs will only come with entirely different learning mechanisms. Humans use much more powerful and efficient ways of learning that "have not yet been properly invented and scaled." This view places him among a growing group of LLM skeptics who argue that the next leap in AI requires new approaches. For example, he proposes "system prompt learning," in which learning takes place at the token and context level rather than by changing the model's weights. He compares it to what happens in the human brain during sleep, when information is consolidated and stored.
Interactive Environments as the Way Forward
One of Karpathy's key proposals is to train LLMs in interactive environments—digital spaces where models can act and observe the consequences of their actions. Earlier phases of training rely on internet text for pretraining and question-and-answer data for fine-tuning, but interactive environments provide genuine feedback based on real decisions. LLMs would thus stop merely imitating human responses statistically and begin learning to make decisions and test choices in controlled scenarios.
Karpathy emphasizes that these environments would serve both training and evaluation. The main challenge now is to create a large, diverse, high-quality collection of such environments, similar to the text datasets of the past. In August 2024, Karpathy argued that RL could be a breakthrough if it relied on genuinely objective, measurable reward functions. At the time, he criticized standard reinforcement learning from human feedback (RLHF) as being too dependent on human preferences, describing it as more of a "vibe check" than a real objective.
Comparison with the Views of Other Experts
Karpathy's ideas align with calls for a paradigm shift from DeepMind researchers such as Richard Sutton and David Silver in their essay "Welcome to the Era of Experience." Both argue that the next wave of advanced AI cannot merely copy human language or judgments. Instead, AI should become more robust, creative, and adaptable by learning directly from experience and independent actions. Karpathy agrees that current RL techniques are limited when it comes to more abstract reasoning and calls for learning from direct experience instead of imitation.
Such views are becoming increasingly prominent in the AI community, where alternatives to current methods are being sought. For example, reasoning models that rely heavily on RL are driving most of the recent hype around AI, while pretrained models such as GPT-4 show only small gains. Karpathy remains optimistic about the growth of RL fine-tuning but emphasizes the need for innovation to achieve genuine progress.



