Nemotron Super
nvidia/nemotron-superFree, open reasoning engineered for maximum inference efficiency on NVIDIA hardware.
Model overview
Context window
128K
tokens
Input price
—
quoted on request
Output price
—
quoted on request
Weekly volume
—
tokens / week
Open weights
Yes
self-hostable
Capabilities
4
of 11 tags
Description
Nemotron Super is a reasoning model from NVIDIA. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. NVIDIA’s stated focus is open reasoning models tuned for its stack. Its mean it can also be self-hosted under its licence.
Capabilities
- Reasoning
- Spends extra billed tokens thinking before it answers, trading latency and cost for accuracy.
- Tool calling
- Calls functions you define and returns their arguments as structured data, the basis of agents.
- Streaming
- Emits tokens as they are generated, so answers appear progressively rather than all at once.
- Open weights
- The parameters are published, so the model can be inspected, fine-tuned and self-hosted under its licence.
Strengths & weaknesses
Strengths
- Free at point of use
- Reference-grade efficiency
- Open weights
Weaknesses
- Rate limits on free tier
- Ceiling trails paid reasoning models
Pricing
Free| Rate | Price | Unit |
|---|---|---|
| Input | $0.00 | per 1M tokens |
| Output | — not charged | — |
Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.
Context window
128Ktokens
covers prompt and response together, so a long input leaves less room for the answer.
Supported features
| Feature | Support |
|---|---|
| Tool calling | Supported |
| JSON mode | Not supported |
| Streaming | Supported |
| Vision input | Not supported |
| Audio input | Not supported |
| Long context | Not supported |
| Open weights | Supported |
| Reasoning | Supported |
| Fine-tunable | Not supported |
| Batch | Not supported |
| Caching | Not supported |
Example use cases
Prototyping reasoning flows
Free at point of use
Self-hosted inference
Reference-grade efficiency
Education
Open weights
Comparison
| Attribute | Nemotron Super | GLM-4.5 Air | DeepSeek R1 |
|---|---|---|---|
| Context window | 128K | 128K | 164K |
| Input price | $0.00 | $0.00 | $0.00 |
| Output price | — not charged | — not charged | — not charged |
| Tool calling | Supported | Supported | Supported |
| JSON mode | Not supported | Not supported | Not supported |
| Vision input | Not supported | Not supported | Not supported |
| Open weights | Supported | Supported | Supported |
| Prompt caching | Not supported | Not supported | Not supported |
Best value in each row is highlighted. Illustrative placeholder data.
Provider information
View NVIDIA →NVIDIA · Santa Clara, United States · est. 1993 (Nemotron open-model programme from 2024)
NVIDIA's Nemotron programme releases open reasoning models engineered to run superbly on its own hardware, free to use, meticulously optimised, and intended to grow the total market for accelerated inference. One model, deliberately: a reference implementation of efficient open reasoning.
Recent releases
Documentation
Frequently asked questions
Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Every figure on this page is an illustrative placeholder.