Scaling AI Isn’t Everything. Nvidia and a Hong Kong University Unveil a Model Stronger Than Its Rivals.

Scaling AI Isn’t Everything. Nvidia and a Hong Kong University Unveil a Model Stronger Than Its Rivals.

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
2. 12. 2025
1 minutes reading
Scaling AI Isn’t Everything. Nvidia and a Hong Kong University Unveil a Model Stronger Than Its Rivals.

Imagine a small model with 8 billion parameters called Orchestrator, which acts like a conductor in an orchestra full of tools and intelligent models. This creation from researchers at Nvidia and the University of Hong Kong solves complex tasks, such as those from Humanity's Last Exam (HLE), where it achieved a score of 37.1%, while GPT-5 scored only 35.1%. And all that at 2.5 times lower cost! No giant model, just smart coordination.

Why are small models more powerful than large ones?

Orchestrator is no lone hero—it calls on tools such as the Tavily search API for web searches, a Python sandbox for running code, or specialized models such as Qwen2.5-Math-72B for mathematics. During training, it uses reinforcement learning with rewards for correct results, low costs, and adherence to user preferences. For example, on the FRAMES benchmark, it outperformed GPT-5 with a success rate of 76.3% at just 30% of the cost.

Benchmark results and cost
Benchmark results and cost

In each round, Orchestrator reasons, selects a tool—such as GPT-5-mini for coding or Llama-3.3-70B-Instruct for more general tasks—and then processes the response. The researchers created the ToolScale dataset with thousands of examples from fields such as finance, sports, and medicine, where the model learns to coordinate up to 50 rounds of interactions. The result? On τ²-Bench, it achieved 80.2%, while calling GPT-5 in only 40% of cases, yet still performed better than GPT-5 alone.

Customization for everyone

Users can set their preferences—for example, prioritizing local search over internet search for privacy reasons. Orchestrator adapts accordingly, making it flexible even with unfamiliar tools such as Claude Opus 4.1 or DeepSeekMath-7b-Instruct. The entire system is designed to be fast and inexpensive, with latency measured in minutes and costs in cents. You can read the detailed report at arxiv.org.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok