Embed v4
cohere/embed-v4Multimodal enterprise embeddings: text and images in one vector space, 128K inputs.
Model overview
Context window
128K
tokens
Input price
—
quoted on request
Output price
—
quoted on request
Weekly volume
—
tokens / week
Open weights
No
API only
Capabilities
2
of 11 tags
Description
Embed v4 is an embeddings model from Cohere. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. Cohere’s stated focus is enterprise retrieval and language stack. The weights are not published, so it is API-only.
Capabilities
- Batch
- Accepts offline batch jobs, which are processed asynchronously at a reduced rate.
- Vision input
- Accepts images alongside text in the prompt.
Strengths & weaknesses
Strengths
- Text + image in one space
- 128K input context
- Enterprise-grade multilingual recall
Weaknesses
- Higher price than Qwen3 Embedding
Pricing
Free| Rate | Price | Unit |
|---|---|---|
| Input | $0.00 | per 1M tokens |
| Output | — not charged | — |
Batch jobs are supported and settle at a reduced rate against the same unit. Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.
Context window
128Ktokens
covers prompt and response together, so a long input leaves less room for the answer.
Supported features
| Feature | Support |
|---|---|
| Tool calling | Not supported |
| JSON mode | Not supported |
| Streaming | Not supported |
| Vision input | Supported |
| Audio input | Not supported |
| Long context | Not supported |
| Open weights | Not supported |
| Reasoning | Not supported |
| Fine-tunable | Not supported |
| Batch | Supported |
| Caching | Not supported |
Example use cases
Multimodal search
Text + image in one space
Enterprise RAG
128K input context
Catalogue matching
Enterprise-grade multilingual recall
Comparison
| Attribute | Embed v4 | Qwen3 Embedding | Gemini Embedding |
|---|---|---|---|
| Context window | 128K | 32K | 8K |
| Input price | $0.00 | $0.00 | $0.00 |
| Output price | — not charged | — not charged | — not charged |
| Tool calling | Not supported | Not supported | Not supported |
| JSON mode | Not supported | Not supported | Not supported |
| Vision input | Supported | Not supported | Not supported |
| Open weights | Not supported | Not supported | Not supported |
| Prompt caching | Not supported | Not supported | Not supported |
Best value in each row is highlighted. Illustrative placeholder data.
Provider information
View Cohere →Cohere · Toronto, Canada · est. 2019
Cohere never chased consumer chat. It built the enterprise retrieval stack instead.
Recent releases
Documentation
Frequently asked questions
Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Batch jobs settle at a reduced rate against the same unit. Every figure on this page is an illustrative placeholder.