Olmo 3.1: Longer Training Boosts AI Without Major Changes

Olmo 3.1: Longer Training Boosts AI Without Major Changes

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
17. 12. 2025
2 minutes reading · 5 views
Olmo 3.1: Longer Training Boosts AI Without Major Changes

The Allen Institute for AI, known as Ai2, has unveiled its most powerful model family to date, called Olmo 3.1. This development comes shortly after Olmo 3 and focuses on efficiency, transparency, and controllability. Instead of devising new architectures, the team simply extended the existing training process using reinforcement learning. This means the model learns from mistakes and successes for longer, making it better at complex tasks such as mathematics or multi-step problem-solving.

Why extend training?

With open AI models, training often stops too early due to high costs. This causes models to lag behind in difficult reasoning or instruction following. Ai2 decided to revisit the question of whether reinforcement learning still helps after the initial gains begin to taper off. The answer is yes. The team took the existing process from Olmo 3 and let it run longer—specifically, for an additional 21 days on 224 graphics processing units (GPUs). This enabled them to advance the model to 32 billion parameters without changing its underlying design. The result? The model improved in mathematics, coding, and complex tasks, proving that scale and patience play a key role.

What exactly was released?

Ai2 updated two versions from Olmo 3: Olmo 3.1 Think 32B, which is optimized for advanced research, and Olmo 3.1 Instruct 32B, designed for instruction following, multi-turn conversations and tool use. A third version, Olmo 3-Base, remains available for programming, comprehension, and mathematics, making it suitable for further fine-tuning. In addition, the team improved the RL-Zero 7B models for mathematics and coding, which serve as baseline references for research. All these models are fully open, meaning anyone can view the data, code, and training decisions.

Performance improvements and availability

The new models delivered noticeable gains in tests. Olmo 3.1 Think 32B scored 5 points higher on AIME, 4 points higher on ZebraLogic, 4 points higher on IFEval, and 20 points higher on IFBench compared to Olmo 3 Think 32B. It also improved in coding and complex multi-step reasoning. In the AIME 2025 benchmark, it outperformed the Qwen 3 32B models and came close to Gemma 27B. Olmo 3.1 Instruct 32B, meanwhile, excels in chat, tool use, and dialogue, where it outperformed its open competitors, including Gemma 3 in mathematics. According to Ai2, Olmo 3.1 Instruct 32B is its most capable fully open chat model at the 32-billion-parameter scale. The RL-Zero 7B models for mathematics and coding benefited from longer and more stable training.

The model weights and checkpoints are available for download on Hugging Face. You can test them in the Ai2 Playground. Released datasets and training artifacts are available for further modifications, enabling fine-tuning or extended reinforcement learning. API access is also coming soon. Ai2 emphasizes transparency—companies can add their own data and retrain the model so that it learns from new information. This is part of Ai2's long-term commitment to open source, including the OlmoTrace tool, which tracks how model outputs correspond to training data.

Source: venturebeat.com

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok