Olmo 3.1: Longer Training Boosts AI Without Major Changes

Olmo 3.1: Longer Training Boosts AI Without Major Changes

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
17. 12. 2025
2 minutes reading
Olmo 3.1: Longer Training Boosts AI Without Major Changes

The Allen Institute for AI, known as Ai2, has unveiled its most powerful model family to date, called Olmo 3.1. This development comes shortly after Olmo 3 and focuses on efficiency, transparency, and controllability. Instead of devising new architectures, the team simply extended the existing training process using reinforcement learning. This means the model learns from mistakes and successes for longer, making it better at complex tasks such as mathematics or multi-step problem-solving.

Why extend training?

With open AI models, training often stops too early due to high costs. This causes models to lag behind in difficult reasoning or instruction following. Ai2 decided to revisit the question of whether reinforcement learning still helps after the initial gains begin to taper off. The answer is yes. The team took the existing process from Olmo 3 and let it run longer—specifically, for an additional 21 days on 224 graphics processing units (GPUs). This enabled them to advance the model to 32 billion parameters without changing its underlying design. The result? The model improved in mathematics, coding, and complex tasks, proving that scale and patience play a key role.

What exactly was released?

Ai2 updated two versions from Olmo 3: Olmo 3.1 Think 32B, which is optimized for advanced research, and Olmo 3.1 Instruct 32B, designed for instruction following, multi-turn conversations and tool use. A third version, Olmo 3-Base, remains available for programming, comprehension, and mathematics, making it suitable for further fine-tuning. In addition, the team improved the RL-Zero 7B models for mathematics and coding, which serve as baseline references for research. All these models are fully open, meaning anyone can view the data, code, and training decisions.

Performance improvements and availability

The new models delivered noticeable gains in tests. Olmo 3.1 Think 32B scored 5 points higher on AIME, 4 points higher on ZebraLogic, 4 points higher on IFEval, and 20 points higher on IFBench compared to Olmo 3 Think 32B. It also improved in coding and complex multi-step reasoning. In the AIME 2025 benchmark, it outperformed the Qwen 3 32B models and came close to Gemma 27B. Olmo 3.1 Instruct 32B, meanwhile, excels in chat, tool use, and dialogue, where it outperformed its open competitors, including Gemma 3 in mathematics. According to Ai2, Olmo 3.1 Instruct 32B is its most capable fully open chat model at the 32-billion-parameter scale. The RL-Zero 7B models for mathematics and coding benefited from longer and more stable training.

The model weights and checkpoints are available for download on Hugging Face. You can test them in the Ai2 Playground. Released datasets and training artifacts are available for further modifications, enabling fine-tuning or extended reinforcement learning. API access is also coming soon. Ai2 emphasizes transparency—companies can add their own data and retrain the model so that it learns from new information. This is part of Ai2's long-term commitment to open source, including the OlmoTrace tool, which tracks how model outputs correspond to training data.

Source: venturebeat.com

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok