Skip to content

Grok 4 Fast

xai/grok-4-fast
Language

Two million tokens of context at commodity prices, long-context economics, reset.

Model overview

Context window

2M

tokens

Input price

quoted on request

Output price

quoted on request

Weekly volume

tokens / week

Open weights

No

API only

Capabilities

4

of 11 tags

Description

Grok 4 Fast is a language model from xAI. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. xAI’s stated focus is real-time-grounded frontier models. The weights are not published, so it is API-only.

Capabilities

Tool calling
Calls functions you define and returns their arguments as structured data, the basis of agents.
JSON mode
Constrains the response to valid JSON matching a schema you supply.
Streaming
Emits tokens as they are generated, so answers appear progressively rather than all at once.
Long context
Built for inputs well beyond the mainstream window, whole repositories, archives or corpora.

Strengths & weaknesses

Strengths

  • 2M context at workhorse prices
  • Fast
  • Good tool calling

Weaknesses

  • Depth trails Grok 4.1
  • Recall softens at extreme depth

Pricing

Free
Pricing for Grok 4 Fast, per 1M tokens
RatePriceUnit
Input$0.00per 1M tokens
Output not charged

Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.

Context window

2Mtokens

covers prompt and response together, so a long input leaves less room for the answer.

2M tokens against a peer range of 32K to 10M, median 256K. Compared across 24 language models. Logarithmic scale.

Supported features

Feature support for Grok 4 Fast
FeatureSupport
Tool callingSupported
JSON modeSupported
StreamingSupported
Vision inputNot supported
Audio inputNot supported
Long contextSupported
Open weightsNot supported
ReasoningNot supported
Fine-tunableNot supported
BatchNot supported
CachingNot supported

Example use cases

  • Long-context volume work

    2M context at workhorse prices

  • Archive Q&A

    Fast

  • Cost-tiered routing

    Good tool calling

Comparison

Grok 4 Fast compared with Llama 4 Scout and MiniMax Text 01
AttributeGrok 4 FastLlama 4 ScoutMiniMax Text 01
Context window2M10M1M
Input price$0.00$0.00$0.00
Output price not charged not charged not charged
Tool callingSupportedSupportedSupported
JSON modeSupportedNot supportedNot supported
Vision inputNot supportedNot supportedNot supported
Open weightsNot supportedSupportedSupported
Prompt cachingNot supportedNot supportedNot supported

Best value in each row is highlighted. Illustrative placeholder data.

Provider information

View xAI
xAI2 models

xAI · San Francisco, United States · est. 2023

xAI built one of the largest training clusters in existence and iterated Grok to the frontier at unusual speed. Its models are distinguished by real-time grounding and enormous fast-inference context, Grok 4 Fast's 2M-token window at commodity prices reset long-context economics.

Recent releases

    Documentation

    Frequently asked questions

    Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Every figure on this page is an illustrative placeholder.