The Allen Institute for AI, known as Ai2, has unveiled its most powerful model family to date, called Olmo 3.1. This development comes shortly after Olmo 3 and focuses on efficiency, transparency, and controllability. Instead of devising new architectures, the team simply extended the existing training process using reinforcement learning. This means the model learns from mistakes and successes for longer, making it better at complex tasks such as mathematics or multi-step problem-solving.
Why extend training?
With open AI models, training often stops too early due to high costs. This causes models to lag behind in difficult reasoning or instruction following. Ai2 decided to revisit the question of whether reinforcement learning still helps after the initial gains begin to taper off. The answer is yes. The team took the existing process from Olmo 3 and let it run longer—specifically, for an additional 21 days on 224 graphics processing units (GPUs). This enabled them to advance the model to 32 billion parameters without changing its underlying design. The result? The model improved in mathematics, coding, and complex tasks, proving that scale and patience play a key role.
What exactly was released?
Ai2 updated two versions from Olmo 3: Olmo 3.1 Think 32B, which is optimized for advanced research, and Olmo 3.1 Instruct 32B, designed for instruction following, multi-turn conversations and tool use. A third version, Olmo 3-Base, remains available for programming, comprehension, and mathematics, making it suitable for further fine-tuning. In addition, the team improved the RL-Zero 7B models for mathematics and coding, which serve as baseline references for research. All these models are fully open, meaning anyone can view the data, code, and training decisions.
Performance improvements and availability
The new models delivered noticeable gains in tests. Olmo 3.1 Think 32B scored 5 points higher on AIME, 4 points higher on ZebraLogic, 4 points higher on IFEval, and 20 points higher on IFBench compared to Olmo 3 Think 32B. It also improved in coding and complex multi-step reasoning. In the AIME 2025 benchmark, it outperformed the Qwen 3 32B models and came close to Gemma 27B. Olmo 3.1 Instruct 32B, meanwhile, excels in chat, tool use, and dialogue, where it outperformed its open competitors, including Gemma 3 in mathematics. According to Ai2, Olmo 3.1 Instruct 32B is its most capable fully open chat model at the 32-billion-parameter scale. The RL-Zero 7B models for mathematics and coding benefited from longer and more stable training.
The model weights and checkpoints are available for download on Hugging Face. You can test them in the Ai2 Playground. Released datasets and training artifacts are available for further modifications, enabling fine-tuning or extended reinforcement learning. API access is also coming soon. Ai2 emphasizes transparency—companies can add their own data and retrain the model so that it learns from new information. This is part of Ai2's long-term commitment to open source, including the OlmoTrace tool, which tracks how model outputs correspond to training data.
Source: venturebeat.com



