Scaling AI Isn’t Everything. Nvidia and a Hong Kong University Unveil a Model Stronger Than Its Rivals.

Scaling AI Isn’t Everything. Nvidia and a Hong Kong University Unveil a Model Stronger Than Its Rivals.

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
2. 12. 2025
1 minutes reading · 4 views
Scaling AI Isn’t Everything. Nvidia and a Hong Kong University Unveil a Model Stronger Than Its Rivals.

Imagine a small model with 8 billion parameters called Orchestrator, which acts like a conductor in an orchestra full of tools and intelligent models. This creation from researchers at Nvidia and the University of Hong Kong solves complex tasks, such as those from Humanity's Last Exam (HLE), where it achieved a score of 37.1%, while GPT-5 scored only 35.1%. And all that at 2.5 times lower cost! No giant model, just smart coordination.

Why are small models more powerful than large ones?

Orchestrator is no lone hero—it calls on tools such as the Tavily search API for web searches, a Python sandbox for running code, or specialized models such as Qwen2.5-Math-72B for mathematics. During training, it uses reinforcement learning with rewards for correct results, low costs, and adherence to user preferences. For example, on the FRAMES benchmark, it outperformed GPT-5 with a success rate of 76.3% at just 30% of the cost.

Benchmark results and cost
Benchmark results and cost

In each round, Orchestrator reasons, selects a tool—such as GPT-5-mini for coding or Llama-3.3-70B-Instruct for more general tasks—and then processes the response. The researchers created the ToolScale dataset with thousands of examples from fields such as finance, sports, and medicine, where the model learns to coordinate up to 50 rounds of interactions. The result? On τ²-Bench, it achieved 80.2%, while calling GPT-5 in only 40% of cases, yet still performed better than GPT-5 alone.

Customization for everyone

Users can set their preferences—for example, prioritizing local search over internet search for privacy reasons. Orchestrator adapts accordingly, making it flexible even with unfamiliar tools such as Claude Opus 4.1 or DeepSeekMath-7b-Instruct. The entire system is designed to be fast and inexpensive, with latency measured in minutes and costs in cents. You can read the detailed report at arxiv.org.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok