Skip to content

Step-1V

stepfun/step-1v
Vision

Vision-first model from a lab that trains multimodality natively rather than bolting it on.

Model overview

Context window

32K

tokens

Input price

quoted on request

Output price

quoted on request

Weekly volume

tokens / week

Open weights

No

API only

Capabilities

2

of 11 tags

Description

Step-1V is a vision model from StepFun. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. StepFun’s stated focus is multimodal foundation models. The weights are not published, so it is API-only.

Capabilities

Vision input
Accepts images alongside text in the prompt.
Streaming
Emits tokens as they are generated, so answers appear progressively rather than all at once.

Strengths & weaknesses

Strengths

  • Natively multimodal training
  • Good spatial reasoning
  • Stable outputs

Weaknesses

  • 32K context
  • Slower ecosystem adoption

Pricing

Free
Pricing for Step-1V, per 1M tokens
RatePriceUnit
Input$0.00per 1M tokens
Output not charged

Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.

Context window

32Ktokens

covers prompt and response together, so a long input leaves less room for the answer.

32K tokens against a peer range of 32K to 128K, median 128K. Compared across 5 vision models.

Supported features

Feature support for Step-1V
FeatureSupport
Tool callingNot supported
JSON modeNot supported
StreamingSupported
Vision inputSupported
Audio inputNot supported
Long contextNot supported
Open weightsNot supported
ReasoningNot supported
Fine-tunableNot supported
BatchNot supported
CachingNot supported

Example use cases

  • Visual reasoning

    Natively multimodal training

  • Scene understanding

    Good spatial reasoning

  • Multimodal research

    Stable outputs

Comparison

Step-1V compared with Step-2 and Hunyuan Vision
AttributeStep-1VStep-2Hunyuan Vision
Context window32K128K128K
Input price$0.00$0.00$0.00
Output price not charged not charged not charged
Tool callingNot supportedSupportedNot supported
JSON modeNot supportedSupportedNot supported
Vision inputSupportedNot supportedSupported
Open weightsNot supportedNot supportedNot supported
Prompt cachingNot supportedNot supportedNot supported

Best value in each row is highlighted. Illustrative placeholder data.

Provider information

View StepFun
StepFun2 models

StepFun · 阶跃星辰 · Shanghai, China · est. 2023

StepFun builds trillion-parameter-class multimodal systems, its Step series spans text and vision with an emphasis on unified multimodal training rather than bolted-on adapters. A quieter lab than its peers, it is consistently cited in multimodal research.

Recent releases

    Documentation

    Frequently asked questions

    Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Every figure on this page is an illustrative placeholder.