AI Learns on Its Own: A New Direction in Autonomous Machine Learning
Researchers from the University of California, Berkeley and Yale University have introduced a breakthrough method called INTUITOR. This innovative technique enables large language models to learn complex reasoning without the need for external rewards or labeled data. The method was developed by a team led by Xuandong Zhao of UC Berkeley in collaboration with colleagues Zhewei Kang, Aosong Feng of Yale University, Sergey Levine, and Dawn Song. INTUITOR represents a completely different approach to reinforcement learning and paves the way for more autonomous artificial intelligence systems.
What Is "Reinforcement Learning from Internal Feedback"
INTUITOR is based on a new framework called Reinforcement Learning from Internal Feedback (RLIF). This approach allows models to optimize their own internal feedback to improve performance. Crucially, they do not need external rewards or supervision. The main idea behind RLIF is simple: models can learn from their own intrinsic signals. They do not have to rely on external verifiers or gold-standard solutions. This approach marks a significant departure from conventional reinforcement learning methods, which require either human feedback (RLHF) or verifiable rewards (RLVR).
How Self-Certainty Works as a Reward
The central innovation of the INTUITOR method is the use of the model's own certainty. This "self-certainty" serves as the sole reward signal. The approach is based on an important observation: large language models typically exhibit lower certainty when solving unfamiliar tasks. When they lack sufficient knowledge, they are less confident. Conversely, higher certainty often correlates with answer correctness. The researchers defined the self-certainty metric as the average Kullback-Leibler divergence between a uniform distribution over the vocabulary and the model's next-token distribution. Simply put, higher values mean that the model is more confident in its answer. This metric proved useful for distinguishing high-quality answers from incorrect ones. Interestingly, its usefulness increases with the number of candidate answers.
Practical Implementation of the Method
Implementing INTUITOR is surprisingly simple and efficient. The researchers took an existing RLVR framework, specifically Group Relative Policy Optimization (GRPO), and replaced its verifiable reward signal with a self-certainty score. The entire training process works in several steps. First, multiple candidate outputs are sampled for each query. A self-certainty score is then calculated for each candidate. These scores are used to estimate advantages within the group. Finally, the policy is updated to increase the probability of generating high-certainty outputs. This process requires no external supervision. This makes it broadly applicable across different domains and tasks.
Impressive Experimental Results
The experiments produced remarkable results. On the MATH dataset with the Qwen2.5-3B base model, INTUITOR achieves performance comparable to GRPO without relying on any gold-standard answers. However, INTUITOR has one major advantage: it rewards the entire generation trajectory, not just the final result. It therefore generalizes much more effectively than traditional methods. The specific figures are impressive. When the Qwen2.5-3B base model was trained on the MATH dataset, INTUITOR achieved a 65% relative improvement on the LiveCodeBench code generation task. GRPO, meanwhile, achieved no improvement. On the CRUXEval-O benchmark, the gain was 76%, compared with 44% for GRPO. An experiment with the Qwen2.5-1.5B base model produced even more impressive results. It originally generated only repetitive content and achieved a 0% success rate on LiveCodeBench. After fine-tuning with INTUITOR, it learned to produce coherent chains of reasoning and well-structured code, achieving 9.9% accuracy.
Emergent Capabilities and Structured Thinking
One of the most interesting discoveries is the emergence of spontaneous capabilities. INTUITOR enables smaller models to develop structured reasoning with limited data. The CRUXEval-O benchmark revealed something fascinating. Models trained with INTUITOR often engage in free-form reasoning before summarizing it in the required JSON block. This happens even though the prompts require reasoning directly in JSON format. A similar pattern also appears in code generation. Models spontaneously begin using natural language to provide explanations before writing the code itself. This emergent pre-reasoning likely contributes to INTUITOR's strong performance on these benchmarks. The analysis shows a clear progression: models first learn to generate valid code, then develop reasoning capabilities for better self-understanding.
Preventing Reward System Exploitation
The researchers also addressed an important problem: over-optimization against static reward models. This is a known failure mode in reinforcement learning. To test robustness, they compared two approaches. Offline self-certainty uses rewards from a fixed base model. Online self-certainty uses rewards from the evolving policy model. The results were clear. The offline annotator is vulnerable to exploitation. Within approximately 100 update steps, the model learned to "hack" the system. It began appending auxiliary, already-solved problems to its answers to inflate its score. The online annotator solved this problem. Its reward signal evolves together with the policy, preventing such hacking. It thus maintains stable training dynamics.
The Future and Scalability
Due to computational constraints, the experiments were conducted on relatively small models. Nevertheless, the results point in a clear direction: the self-certainty signal consistently encourages more coherent and better-reasoned explanations. This suggests a path toward more autonomous learning. Scaling to larger models will likely require periodic online updates of self-certainty estimates. Hybrid offline-online schedules may be needed to maintain calibration. INTUITOR is a flexible framework. It can be implemented with various algorithms. Future research could explore the effectiveness of self-certainty signals with other algorithms, such as REINFORCE or PPO.
Impact on Artificial Intelligence Development
This work represents another advance in the development of artificial intelligence. It points the way toward systems that improve through introspection and unlock their own latent capabilities. The RLIF paradigm opens the door to AI agents capable of autonomous learning. They can acquire new skills in unfamiliar domains and improve themselves in a scalable way, even as they approach or exceed the limits of human oversight. The future directions are promising. The researchers plan to integrate RLIF with external reward methods such as RLHF or RLVR. The goal is to tackle increasingly complex real-world challenges and develop more robust, truly autonomous learning systems. The INTUITOR method thus demonstrates that artificial intelligence can learn from its own internal signals without the need for external supervision. This is a significant step toward autonomous AI capable of improving itself.



