Skip to content

DeepSeek V4 Flash

deepseek/deepseek-v4-flash
Language

Very low cost per token, with peak-hour pricing at 2x.

Model overview

Context window

1.048576M

tokens

Input price

R4.60

per 1M tokens

Output price

R13.70

per 1M tokens

Weekly volume

tokens / week

Open weights

No

API only

Capabilities

3

of 11 tags

Description

DeepSeek V4 Flash is a language model from DeepSeek. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. DeepSeek’s stated focus is open-weight frontier reasoning at disruptive economics. The weights are not published, so it is API-only.

Capabilities

Tool calling
Calls functions you define and returns their arguments as structured data, the basis of agents.
JSON mode
Constrains the response to valid JSON matching a schema you supply.
Streaming
Emits tokens as they are generated, so answers appear progressively rather than all at once.

Strengths & weaknesses

Strengths

  • Cheapest per token on the rate card
  • Open weights permit self-hosting later
  • Price quoted from the supplier's own rate card, not estimated

Weaknesses

  • Priced at 2x during the supplier's peak hours
  • English long-form prose trails Western flagships
  • Not provisioned on this account yet — lead time applies

Pricing

Pricing for DeepSeek V4 Flash, per 1M tokens
RatePriceUnit
InputR4.60per 1M tokens
OutputR13.70per 1M tokens

This model is not provisioned yet; the rate is quoted on request.

Context window

1.048576Mtokens

covers prompt and response together, so a long input leaves less room for the answer.

1.048576M tokens against a peer range of 32K to 10M, median 256K. Compared across 34 language models. Logarithmic scale.

Supported features

Feature support for DeepSeek V4 Flash
FeatureSupport
Tool callingSupported
JSON modeSupported
StreamingSupported
Vision inputNot supported
Audio inputNot supported
Long contextNot supported
Open weightsNot supported
ReasoningNot supported
Fine-tunableNot supported
BatchNot supported
CachingNot supported

Example use cases

  • General assistants

  • Retrieval-augmented answering

  • Everyday production traffic

DeepSeek V4 Flash compared with DeepSeek V4 Pro and AC Flash 4.5
AttributeDeepSeek V4 FlashDeepSeek V4 ProAC Flash 4.5
Context window1.048576M1.048576M
Input priceR4.60R13.70R10.00
Output priceR13.70R41.30R40.00
Tool callingSupportedSupportedSupported
JSON modeSupportedSupportedSupported
Vision inputNot supportedNot supportedNot supported
Open weightsNot supportedNot supportedNot supported
Prompt cachingNot supportedNot supportedNot supported

Best value in each row is highlighted. Blank cells are facts the supplier does not publish, not zeroes.

Provider information

View DeepSeek
DeepSeek5 models

DeepSeek · 深度求索 · Hangzhou, China · est. 2023

Spun out of a quantitative trading firm, DeepSeek built its reputation by releasing open-weight models that matched closed frontier quality at a fraction of the price. Its V-series redefined cost expectations for production LLMs, and its R-series brought open reasoning models into serious enterprise conversations.

Recent releases

    Documentation

    Frequently asked questions

    Input is billed at R4.60 per 1M tokens and output at R13.70 per 1M tokens. Every figure on this page is an illustrative placeholder.