Beyond GPUs, NVIDIA ships the Nemotron family of open-weight models — designed to be post-trained on your own data for enterprise agents, legal, search, clinical and sovereign-AI workloads at a fraction of closed-frontier cost.
Nemotron 3 Ultra / Nano Omni are the current open line; weights are free to self-host.
Who they are
NVIDIA’s Nemotron Labs pitches open models as the customization layer for enterprises: keep full control of weights, inspect training, run private evals and stand up agents on your own infrastructure. Deployments include Harvey (legal), Glean (agentic search, Waldo), H Company (computer use) and YTL AI Labs (Malay sovereign AI).
Model lineup
Model
What it is for
Vendor list price
Nemotron 3 Ultraopen flagship
Frontier-class open model for enterprise post-training; top open-model agent accuracy via LangChain Deep Agents; Harvey matched legal frontier at ~10x lower cost
Open weights — free to self-host
Nemotron 3 Nano Omnismall open multimodal
Small-model computer-use: H Company’s Holotron 3 Nano scored over 76% on OSWorld-Verified
Open weights — free to self-host
Nemotron 3.5 Lightning30B MoE / 3B active · 11 Aug 2026
Resident-agent workhorse — ~4x faster inference, paired with the open NeMo Switchyard routing library
Open weights — free to self-host
Last updated — Nemotron is open-weight: the model licence is free for self-hosting. Because there is no vendor list price, the cost that matters is whatever your host charges — a post-trained deployment’s per-token cost depends on the hardware and the batch size, and the third-party figures in circulation vary too widely to quote as a single rate. Hosted enterprise access runs via NVIDIA NIM / build.nvidia; self-hosted cost is your own GPU time. Figures are vendor public list prices in USD per 1M tokens, standard tier — no batch, cache, promo or peak-hour discount applied. Vendors change list prices without notice: confirm on the vendor’s own pricing page before budgeting.
Historical models — for reference
Discontinued or fully superseded models, kept for historical reference. Active models — even previous-generation ones still on sale — always stay in the lineup above.
Model
Era & what it was for
Historic list price
Llama-3.3-Nemotron-Super-49BMar 2025 · open distillation
The 2025 open enterprise workhorse distilled from Llama 3.3 70B — reasoning and tool-use at a single-GPU-friendly size
Open weights — free to self-host
Llama-3.1-Nemotron-70B2024 · NVIDIA-refined Llama
NVIDIA’s first widely adopted refined-Llama open model — strong instruction following for its era
Open weights — free to self-host
Latest news
Recent announcements with outlet links (EN first, CN where available).
14 Jul 2026
NVIDIA pitches Nemotron as the enterprise tuning layer
Nemotron Labs listed six production deployments — Harvey, Glean, H Company, Abridge, Heidi Health and YTL AI Labs — matching closed-frontier accuracy at 10–20x lower cost per run.