N

Nemotron 3 Ultra

LLMby NVIDIA·Model page

NVIDIA's 550B-parameter MoE LLM with a 1M-token context for enterprise-scale text generation.

Max output

Most tokens the model can return in a single response.

33Ktokens
Share:

Design Arena

Design Arena ranks models on real-world front-end and design tasks — websites, UI components, data viz, SVG and more — through head-to-head human votes, scored as an ELO rating.
CategoryELOWin rateRank
3D116641.0%#58
Data viz115738.3%#70
UI component115737.3%#70
Code115536.3%#76
Game dev115436.9%#69

Pricing

per 1M tokens
Input$0.625 /1M
Output$3.13 /1M
Cache read$0.188 /1M

Model Description

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Author
N
NVIDIA
Organization · ✓
nvidia
Details
Downloads
Likes
AccessOpen Source
Context262K tokens
Input price$0.625 /1M
Output price$3.13 /1M
CreatedJun 4, 2026
Updated
View on Hugging Face
Benchmarks
Coding49.3
Agentic21.7
Get the full context.

Sign up to read complete case studies, access detailed metrics, and unlock all use cases.

Nemotron 3 Ultra — AI Model Details | Applied