NVIDIA Unveils Open Nemotron 3 AI Models

NVIDIA Unveils Open Nemotron 3 AI Models

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
19. 12. 2025
4 minutes reading
NVIDIA Unveils Open Nemotron 3 AI Models

Artificial intelligence (AI) works like a team of smart assistants collaborating on complex tasks. This is exactly what the new Nvidia Nemotron 3 model family enables. This suite of open models comes in three sizes: Nano, Super, and Ultra. It is designed for creating specialized agentic AI systems. Nvidia announced it on December 15, 2025, with a focus on high efficiency and accuracy, helping developers build systems that can handle long contexts and many agents simultaneously.

The Nemotron 3 family addresses problems such as high computational demands and loss of context during lengthy tasks. The models are open, meaning developers can inspect how they work, modify them, and deploy them anywhere. Jensen Huang, founder and CEO of Nvidia, emphasized that open innovation is the foundation of progress in AI. The models support sovereign AI initiatives in European countries and South Korea, where companies want systems tailored to their data and regulations.

What Nemotron 3 Nano offers

Nemotron 3 Nano is the first model in the family to be available immediately. It has a total of 30 billion parameters but activates only up to 3 billion at a time, making it ideal for tasks such as software debugging, content summarization, and information retrieval. It uses a hybrid mixture-of-experts (MoE) architecture that combines Mamba layers for fast sequence processing with Transformer layers for precise reasoning. This combination delivers up to 4x higher token throughput than the previous Nemotron 2 Nano and reduces the generation of reasoning tokens by up to 60%.

The model handles a context of up to 1 million tokens, meaning it can retain a vast amount of information without losing context. According to evaluations by Artificial Analysis, it achieves the highest accuracy among similarly sized models, with an intelligence index of 52. Developers can run it on GPUs such as DGX Spark, H100, or B200. Nvidia provides ready-made guides for inference engines such as vLLM, SGLang, and TRT-LLM, enabling rapid deployment.

Advanced technologies in Nemotron 3

All Nemotron 3 models use multi-environment reinforcement learning (RL) through the open NeMo Gym library. This method trains the model on sequences of actions across different environments, improving reliability in multi-step tasks such as tool calling or writing functional code. NeMo Gym is open, allowing developers to create custom environments for specific domains.

The hybrid MoE architecture in Nemotron 3 interleaves Mamba-2 and MoE layers with several self-attention layers, maximizing inference speed while maintaining accuracy. Mamba layers process long-range dependencies with low memory usage, while Transformer layers handle logical relationships in tasks such as mathematics or planning. MoE activates only a subset of experts for each token, reducing latency and increasing throughput, making it ideal for systems with many agents.

For long contexts, the model supports 1 million tokens, allowing it to work with large codebases or documents without splitting them into chunks. This improves factual accuracy in applications such as retrieval-augmented generation or compliance analysis.

What is included in the Nemotron 3 Super and Ultra models

Nemotron 3 Super and Ultra will arrive in the first half of 2026. Super has approximately 100 billion parameters, with up to 10 billion active per token, making it suitable for systems with many collaborating agents. Ultra has around 500 billion parameters, with up to 50 billion active, and is intended for complex tasks such as deep research.

These models introduce latent MoE, where experts operate on a shared latent representation, enabling 4x more experts at the same inference cost. In addition, multi-token prediction increases speed by 2.4% during training and enables speculative decoding. Training is performed in NVFP4, Nvidia's 4-bit format, which reduces memory requirements and accelerates the process on the Blackwell architecture while maintaining accuracy.

The models are trained on 25 trillion tokens, and Nvidia is releasing nearly 10 trillion tokens of synthetic corpus data for inspection. The datasets include 3 trillion tokens for pretraining, with extensive coverage of code and mathematics, 13 million samples for post-training, and datasets for RL. The new Nemotron Agentic Safety Dataset contains nearly 11,000 traces of agentic workflows for safety evaluation.

Openness and support for developers

Nvidia is releasing the model weights under the Nvidia Open Model License, along with training recipes in the Nemotron GitHub repository. This includes pretraining recipes, RL alignment, and data pipelines. Developers can customize the models for their needs, such as domain-specific tasks.

Nemotron 3 Nano is available on Hugging Face, Baseten, DeepInfra, Fireworks, FriendliAI, OpenRouter, and Together AI. It is supported by platforms such as Couchbase, DataRobot, H2O.ai, JFrog, Lambda, and UiPath, with support from AWS, Google Cloud, CoreWeave, Crusoe, Microsoft Foundry, Nebius, Nscale, and Yotta coming soon. As an Nvidia NIM microservice, it enables secure deployment.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok