Skip to content

Embed v4

cohere/embed-v4
Embeddings

Multimodal enterprise embeddings: text and images in one vector space, 128K inputs.

Model overview

Context window

128K

tokens

Input price

quoted on request

Output price

quoted on request

Weekly volume

tokens / week

Open weights

No

API only

Capabilities

2

of 11 tags

Description

Embed v4 is an embeddings model from Cohere. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. Cohere’s stated focus is enterprise retrieval and language stack. The weights are not published, so it is API-only.

Capabilities

Batch
Accepts offline batch jobs, which are processed asynchronously at a reduced rate.
Vision input
Accepts images alongside text in the prompt.

Strengths & weaknesses

Strengths

  • Text + image in one space
  • 128K input context
  • Enterprise-grade multilingual recall

Weaknesses

  • Higher price than Qwen3 Embedding

Pricing

Free
Pricing for Embed v4, per 1M tokens
RatePriceUnit
Input$0.00per 1M tokens
Output not charged

Batch jobs are supported and settle at a reduced rate against the same unit. Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.

Context window

128Ktokens

covers prompt and response together, so a long input leaves less room for the answer.

128K tokens against a peer range of 8K to 128K, median 32K. Compared across 3 embeddings models.

Supported features

Feature support for Embed v4
FeatureSupport
Tool callingNot supported
JSON modeNot supported
StreamingNot supported
Vision inputSupported
Audio inputNot supported
Long contextNot supported
Open weightsNot supported
ReasoningNot supported
Fine-tunableNot supported
BatchSupported
CachingNot supported

Example use cases

  • Multimodal search

    Text + image in one space

  • Enterprise RAG

    128K input context

  • Catalogue matching

    Enterprise-grade multilingual recall

Comparison

Embed v4 compared with Qwen3 Embedding and Gemini Embedding
AttributeEmbed v4Qwen3 EmbeddingGemini Embedding
Context window128K32K8K
Input price$0.00$0.00$0.00
Output price not charged not charged not charged
Tool callingNot supportedNot supportedNot supported
JSON modeNot supportedNot supportedNot supported
Vision inputSupportedNot supportedNot supported
Open weightsNot supportedNot supportedNot supported
Prompt cachingNot supportedNot supportedNot supported

Best value in each row is highlighted. Illustrative placeholder data.

Provider information

View Cohere
Cohere3 models

Cohere · Toronto, Canada · est. 2019

Cohere never chased consumer chat. It built the enterprise retrieval stack instead.

Recent releases

    Documentation

    Frequently asked questions

    Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Batch jobs settle at a reduced rate against the same unit. Every figure on this page is an illustrative placeholder.