Skip to content

Gemini 3 Pro

google/gemini-3-pro
Language

Natively multimodal frontier: text, vision and audio in one model with a million-token window.

Model overview

Context window

1M

tokens

Input price

quoted on request

Output price

quoted on request

Weekly volume

tokens / week

Open weights

No

API only

Capabilities

6

of 11 tags

Description

Gemini 3 Pro is a language model from Google. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. Google’s stated focus is frontier multimodality at planetary scale. The weights are not published, so it is API-only.

Capabilities

Tool calling
Calls functions you define and returns their arguments as structured data, the basis of agents.
JSON mode
Constrains the response to valid JSON matching a schema you supply.
Streaming
Emits tokens as they are generated, so answers appear progressively rather than all at once.
Vision input
Accepts images alongside text in the prompt.
Audio input
Accepts audio alongside text in the prompt.
Long context
Built for inputs well beyond the mainstream window, whole repositories, archives or corpora.

Strengths & weaknesses

Strengths

  • True native multimodality including audio
  • 1M context
  • Top multimodal benchmark (MMMU 78)

Weaknesses

  • Output pricing above mid-tier
  • Closed weights

Pricing

Free
Pricing for Gemini 3 Pro, per 1M tokens
RatePriceUnit
Input$0.00per 1M tokens
Output not charged

Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.

Context window

1Mtokens

covers prompt and response together, so a long input leaves less room for the answer.

1M tokens against a peer range of 32K to 10M, median 256K. Compared across 24 language models. Logarithmic scale.

Supported features

Feature support for Gemini 3 Pro
FeatureSupport
Tool callingSupported
JSON modeSupported
StreamingSupported
Vision inputSupported
Audio inputSupported
Long contextSupported
Open weightsNot supported
ReasoningNot supported
Fine-tunableNot supported
BatchNot supported
CachingNot supported

Example use cases

  • Multimodal products

    True native multimodality including audio

  • Long-context analysis

    1M context

  • Audio understanding

    Top multimodal benchmark (MMMU 78)

Comparison

Gemini 3 Pro compared with GPT-5.1 and Claude Sonnet 4.5
AttributeGemini 3 ProGPT-5.1Claude Sonnet 4.5
Context window1M400K1M
Input price$0.00$0.00$0.00
Output price not charged not charged not charged
Tool callingSupportedSupportedSupported
JSON modeSupportedSupportedSupported
Vision inputSupportedSupportedSupported
Open weightsNot supportedNot supportedNot supported
Prompt cachingNot supportedSupportedSupported

Best value in each row is highlighted. Illustrative placeholder data.

Provider information

View Google
Google5 models

Google · Mountain View, United States · est. 1998 (DeepMind model era from 2023)

Google DeepMind's Gemini line is natively multimodal (text, vision and audio in one system) backed by custom TPU infrastructure nobody else can match. Veo and Imagen lead generative media, and Gemini's long-context engineering set the 1M-token standard the industry chased.

Recent releases

    Documentation

    Frequently asked questions

    Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Every figure on this page is an illustrative placeholder.