Skip to content

Qwen VL Max

qwen/qwen-vl-max
Vision

Vision-language specialist strong on documents, charts and OCR-heavy enterprise inputs.

Model overview

Context window

128K

tokens

Input price

quoted on request

Output price

quoted on request

Weekly volume

tokens / week

Open weights

No

API only

Capabilities

3

of 11 tags

Description

Qwen VL Max is a vision model from Alibaba Qwen. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. Alibaba Qwen’s stated focus is the broadest open model family in the world. The weights are not published, so it is API-only.

Capabilities

Vision input
Accepts images alongside text in the prompt.
Tool calling
Calls functions you define and returns their arguments as structured data, the basis of agents.
Streaming
Emits tokens as they are generated, so answers appear progressively rather than all at once.

Strengths & weaknesses

Strengths

  • Excellent document and table understanding
  • Strong CJK OCR
  • Reasonable pricing for vision

Weaknesses

  • Natural-scene reasoning trails document skill
  • No JSON mode

Pricing

Free
Pricing for Qwen VL Max, per 1M tokens
RatePriceUnit
Input$0.00per 1M tokens
Output not charged

Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.

Context window

128Ktokens

covers prompt and response together, so a long input leaves less room for the answer.

128K tokens against a peer range of 32K to 128K, median 128K. Compared across 5 vision models.

Supported features

Feature support for Qwen VL Max
FeatureSupport
Tool callingSupported
JSON modeNot supported
StreamingSupported
Vision inputSupported
Audio inputNot supported
Long contextNot supported
Open weightsNot supported
ReasoningNot supported
Fine-tunableNot supported
BatchNot supported
CachingNot supported

Example use cases

  • Invoice and form extraction

    Excellent document and table understanding

  • Chart reading

    Strong CJK OCR

  • Screenshot QA

    Reasonable pricing for vision

Comparison

Qwen VL Max compared with Qwen3 Max and Kimi VL
AttributeQwen VL MaxQwen3 MaxKimi VL
Context window128K256K128K
Input price$0.00$0.00$0.00
Output price not charged not charged not charged
Tool callingSupportedSupportedNot supported
JSON modeNot supportedSupportedNot supported
Vision inputSupportedSupportedSupported
Open weightsNot supportedNot supportedNot supported
Prompt cachingNot supportedSupportedNot supported

Best value in each row is highlighted. Illustrative placeholder data.

Provider information

View Alibaba Qwen
Alibaba Qwen5 models

Alibaba Qwen · 通义千问 · Hangzhou, China · est. 2023 (Model lab; parent Alibaba founded 1999)

Alibaba's Qwen family spans every size class from edge models to frontier MoE systems, most released with open weights. Its breadth (chat, coding, vision, embeddings) and permissive licensing made it the default base model for much of the global open-source ecosystem.

Recent releases

    Documentation

    Frequently asked questions

    Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Every figure on this page is an illustrative placeholder.