Gemini 3 Pro
google/gemini-3-proNatively multimodal frontier: text, vision and audio in one model with a million-token window.
Model overview
Context window
1M
tokens
Input price
—
quoted on request
Output price
—
quoted on request
Weekly volume
—
tokens / week
Open weights
No
API only
Capabilities
6
of 11 tags
Description
Gemini 3 Pro is a language model from Google. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. Google’s stated focus is frontier multimodality at planetary scale. The weights are not published, so it is API-only.
Capabilities
- Tool calling
- Calls functions you define and returns their arguments as structured data, the basis of agents.
- JSON mode
- Constrains the response to valid JSON matching a schema you supply.
- Streaming
- Emits tokens as they are generated, so answers appear progressively rather than all at once.
- Vision input
- Accepts images alongside text in the prompt.
- Audio input
- Accepts audio alongside text in the prompt.
- Long context
- Built for inputs well beyond the mainstream window, whole repositories, archives or corpora.
Strengths & weaknesses
Strengths
- True native multimodality including audio
- 1M context
- Top multimodal benchmark (MMMU 78)
Weaknesses
- Output pricing above mid-tier
- Closed weights
Pricing
Free| Rate | Price | Unit |
|---|---|---|
| Input | $0.00 | per 1M tokens |
| Output | — not charged | — |
Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.
Context window
1Mtokens
covers prompt and response together, so a long input leaves less room for the answer.
Supported features
| Feature | Support |
|---|---|
| Tool calling | Supported |
| JSON mode | Supported |
| Streaming | Supported |
| Vision input | Supported |
| Audio input | Supported |
| Long context | Supported |
| Open weights | Not supported |
| Reasoning | Not supported |
| Fine-tunable | Not supported |
| Batch | Not supported |
| Caching | Not supported |
Example use cases
Multimodal products
True native multimodality including audio
Long-context analysis
1M context
Audio understanding
Top multimodal benchmark (MMMU 78)
Comparison
| Attribute | Gemini 3 Pro | GPT-5.1 | Claude Sonnet 4.5 |
|---|---|---|---|
| Context window | 1M | 400K | 1M |
| Input price | $0.00 | $0.00 | $0.00 |
| Output price | — not charged | — not charged | — not charged |
| Tool calling | Supported | Supported | Supported |
| JSON mode | Supported | Supported | Supported |
| Vision input | Supported | Supported | Supported |
| Open weights | Not supported | Not supported | Not supported |
| Prompt caching | Not supported | Supported | Supported |
Best value in each row is highlighted. Illustrative placeholder data.
Provider information
View Google →Google · Mountain View, United States · est. 1998 (DeepMind model era from 2023)
Google DeepMind's Gemini line is natively multimodal (text, vision and audio in one system) backed by custom TPU infrastructure nobody else can match. Veo and Imagen lead generative media, and Gemini's long-context engineering set the 1M-token standard the industry chased.
Recent releases
Documentation
Frequently asked questions
Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Every figure on this page is an illustrative placeholder.