Step-1V
stepfun/step-1vVision-first model from a lab that trains multimodality natively rather than bolting it on.
Model overview
Context window
32K
tokens
Input price
—
quoted on request
Output price
—
quoted on request
Weekly volume
—
tokens / week
Open weights
No
API only
Capabilities
2
of 11 tags
Description
Step-1V is a vision model from StepFun. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. StepFun’s stated focus is multimodal foundation models. The weights are not published, so it is API-only.
Capabilities
- Vision input
- Accepts images alongside text in the prompt.
- Streaming
- Emits tokens as they are generated, so answers appear progressively rather than all at once.
Strengths & weaknesses
Strengths
- Natively multimodal training
- Good spatial reasoning
- Stable outputs
Weaknesses
- 32K context
- Slower ecosystem adoption
Pricing
Free| Rate | Price | Unit |
|---|---|---|
| Input | $0.00 | per 1M tokens |
| Output | — not charged | — |
Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.
Context window
32Ktokens
covers prompt and response together, so a long input leaves less room for the answer.
Supported features
| Feature | Support |
|---|---|
| Tool calling | Not supported |
| JSON mode | Not supported |
| Streaming | Supported |
| Vision input | Supported |
| Audio input | Not supported |
| Long context | Not supported |
| Open weights | Not supported |
| Reasoning | Not supported |
| Fine-tunable | Not supported |
| Batch | Not supported |
| Caching | Not supported |
Example use cases
Visual reasoning
Natively multimodal training
Scene understanding
Good spatial reasoning
Multimodal research
Stable outputs
Comparison
| Attribute | Step-1V | Step-2 | Hunyuan Vision |
|---|---|---|---|
| Context window | 32K | 128K | 128K |
| Input price | $0.00 | $0.00 | $0.00 |
| Output price | — not charged | — not charged | — not charged |
| Tool calling | Not supported | Supported | Not supported |
| JSON mode | Not supported | Supported | Not supported |
| Vision input | Supported | Not supported | Supported |
| Open weights | Not supported | Not supported | Not supported |
| Prompt caching | Not supported | Not supported | Not supported |
Best value in each row is highlighted. Illustrative placeholder data.
Provider information
View StepFun →StepFun · 阶跃星辰 · Shanghai, China · est. 2023
StepFun builds trillion-parameter-class multimodal systems, its Step series spans text and vision with an emphasis on unified multimodal training rather than bolted-on adapters. A quieter lab than its peers, it is consistently cited in multimodal research.
Recent releases
Documentation
Frequently asked questions
Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Every figure on this page is an illustrative placeholder.