Skip to content

Gemini Embedding

google/gemini-embedding
Embeddings

High-recall embeddings backed by Gemini's multilingual foundation.

Model overview

Context window

8K

tokens

Input price

quoted on request

Output price

quoted on request

Weekly volume

tokens / week

Open weights

No

API only

Capabilities

1

of 11 tags

Description

Gemini Embedding is an embeddings model from Google. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. Google’s stated focus is frontier multimodality at planetary scale. The weights are not published, so it is API-only.

Capabilities

Batch
Accepts offline batch jobs, which are processed asynchronously at a reduced rate.

Strengths & weaknesses

Strengths

  • Excellent multilingual recall
  • Stable, versioned releases
  • Batch pricing

Weaknesses

  • 8K context per input
  • Text-only

Pricing

Free
Pricing for Gemini Embedding, per 1M tokens
RatePriceUnit
Input$0.00per 1M tokens
Output not charged

Batch jobs are supported and settle at a reduced rate against the same unit. Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.

Context window

8Ktokens

covers prompt and response together, so a long input leaves less room for the answer.

8K tokens against a peer range of 8K to 128K, median 32K. Compared across 3 embeddings models.

Supported features

Feature support for Gemini Embedding
FeatureSupport
Tool callingNot supported
JSON modeNot supported
StreamingNot supported
Vision inputNot supported
Audio inputNot supported
Long contextNot supported
Open weightsNot supported
ReasoningNot supported
Fine-tunableNot supported
BatchSupported
CachingNot supported

Example use cases

  • Semantic search

    Excellent multilingual recall

  • Classification

    Stable, versioned releases

  • Clustering

    Batch pricing

Comparison

Gemini Embedding compared with Embed v4 and Qwen3 Embedding
AttributeGemini EmbeddingEmbed v4Qwen3 Embedding
Context window8K128K32K
Input price$0.00$0.00$0.00
Output price not charged not charged not charged
Tool callingNot supportedNot supportedNot supported
JSON modeNot supportedNot supportedNot supported
Vision inputNot supportedSupportedNot supported
Open weightsNot supportedNot supportedNot supported
Prompt cachingNot supportedNot supportedNot supported

Best value in each row is highlighted. Illustrative placeholder data.

Provider information

View Google
Google5 models

Google · Mountain View, United States · est. 1998 (DeepMind model era from 2023)

Google DeepMind's Gemini line is natively multimodal (text, vision and audio in one system) backed by custom TPU infrastructure nobody else can match. Veo and Imagen lead generative media, and Gemini's long-context engineering set the 1M-token standard the industry chased.

Recent releases

    Documentation

    Frequently asked questions

    Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Batch jobs settle at a reduced rate against the same unit. Every figure on this page is an illustrative placeholder.