Skip to content

Nemotron Super

nvidia/nemotron-super
Open weightsReasoning

Free, open reasoning engineered for maximum inference efficiency on NVIDIA hardware.

Model overview

Context window

128K

tokens

Input price

quoted on request

Output price

quoted on request

Weekly volume

tokens / week

Open weights

Yes

self-hostable

Capabilities

4

of 11 tags

Description

Nemotron Super is a reasoning model from NVIDIA. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. NVIDIA’s stated focus is open reasoning models tuned for its stack. Its mean it can also be self-hosted under its licence.

Capabilities

Reasoning
Spends extra billed tokens thinking before it answers, trading latency and cost for accuracy.
Tool calling
Calls functions you define and returns their arguments as structured data, the basis of agents.
Streaming
Emits tokens as they are generated, so answers appear progressively rather than all at once.
Open weights
The parameters are published, so the model can be inspected, fine-tuned and self-hosted under its licence.

Strengths & weaknesses

Strengths

  • Free at point of use
  • Reference-grade efficiency
  • Open weights

Weaknesses

  • Rate limits on free tier
  • Ceiling trails paid reasoning models

Pricing

Free
Pricing for Nemotron Super, per 1M tokens
RatePriceUnit
Input$0.00per 1M tokens
Output not charged

Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.

Context window

128Ktokens

covers prompt and response together, so a long input leaves less room for the answer.

128K tokens against a peer range of 128K to 300K, median 182K. Compared across 8 reasoning models.

Supported features

Feature support for Nemotron Super
FeatureSupport
Tool callingSupported
JSON modeNot supported
StreamingSupported
Vision inputNot supported
Audio inputNot supported
Long contextNot supported
Open weightsSupported
ReasoningSupported
Fine-tunableNot supported
BatchNot supported
CachingNot supported

Example use cases

  • Prototyping reasoning flows

    Free at point of use

  • Self-hosted inference

    Reference-grade efficiency

  • Education

    Open weights

Comparison

Nemotron Super compared with GLM-4.5 Air and DeepSeek R1
AttributeNemotron SuperGLM-4.5 AirDeepSeek R1
Context window128K128K164K
Input price$0.00$0.00$0.00
Output price not charged not charged not charged
Tool callingSupportedSupportedSupported
JSON modeNot supportedNot supportedNot supported
Vision inputNot supportedNot supportedNot supported
Open weightsSupportedSupportedSupported
Prompt cachingNot supportedNot supportedNot supported

Best value in each row is highlighted. Illustrative placeholder data.

Provider information

View NVIDIA
NVIDIA1 models

NVIDIA · Santa Clara, United States · est. 1993 (Nemotron open-model programme from 2024)

NVIDIA's Nemotron programme releases open reasoning models engineered to run superbly on its own hardware, free to use, meticulously optimised, and intended to grow the total market for accelerated inference. One model, deliberately: a reference implementation of efficient open reasoning.

Recent releases

    Documentation

    Frequently asked questions

    Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Every figure on this page is an illustrative placeholder.