Speech 02
minimax/speech-02Production text-to-speech with expressive multilingual voices, priced per character.
Model overview
Unit price
—
quoted on request
Output price
—
quoted on request
Weekly volume
—
tokens / week
Open weights
No
API only
Capabilities
1
of 11 tags
Description
Speech 02 is a speech model from MiniMax. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. MiniMax’s stated focus is multimodal breadth: text, video, speech. The weights are not published, so it is API-only.
Capabilities
- Streaming
- Emits tokens as they are generated, so answers appear progressively rather than all at once.
Strengths & weaknesses
Strengths
- Expressive, natural voices
- Wide language coverage
- Streaming synthesis
Weaknesses
- Voice cloning requires separate approval
- English slightly behind dedicated Western TTS
Pricing
Free| Rate | Price | Unit |
|---|---|---|
| Unit price | $0.00 | per 1M characters |
Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.
Supported features
| Feature | Support |
|---|---|
| Tool calling | Not supported |
| JSON mode | Not supported |
| Streaming | Supported |
| Vision input | Not supported |
| Audio input | Not supported |
| Long context | Not supported |
| Open weights | Not supported |
| Reasoning | Not supported |
| Fine-tunable | Not supported |
| Batch | Not supported |
| Caching | Not supported |
Example use cases
Voice agents
Expressive, natural voices
Audiobook narration
Wide language coverage
Product voiceovers
Streaming synthesis
Comparison
| Attribute | Speech 02 | Voxtral | Hailuo Video 02 |
|---|---|---|---|
| Input price | $0.00 | $0.00 | $0.00 |
| Output price | — not charged | — not charged | — not charged |
| Tool calling | Not supported | Not supported | Not supported |
| JSON mode | Not supported | Not supported | Not supported |
| Vision input | Not supported | Not supported | Not supported |
| Open weights | Not supported | Supported | Not supported |
| Prompt caching | Not supported | Not supported | Not supported |
Best value in each row is highlighted. Illustrative placeholder data.
Provider information
View MiniMax →MiniMax · 稀宇科技 · Shanghai, China · est. 2021
MiniMax builds across more modalities than almost any peer, million-token text models, the Hailuo video line, and production-grade speech synthesis. Its M-series brought efficient open reasoning to the mix, rounding out one of the most complete catalogues in the ecosystem.
Recent releases
Documentation
Frequently asked questions
Usage is billed at $0.00 per 1M characters. Every figure on this page is an illustrative placeholder.