GLM-4.6
zhipu/glm-4.6The open agentic specialist, engineered for tool loops, computer use and multi-step autonomy.
Model overview
Context window
200K
tokens
Input price
—
quoted on request
Output price
—
quoted on request
Weekly volume
—
tokens / week
Open weights
Yes
self-hostable
Capabilities
5
of 11 tags
Description
GLM-4.6 is an agents model from Zhipu AI. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. Zhipu AI’s stated focus is open agentic GLM family. Its mean it can also be self-hosted under its licence.
Capabilities
- Tool calling
- Calls functions you define and returns their arguments as structured data, the basis of agents.
- JSON mode
- Constrains the response to valid JSON matching a schema you supply.
- Streaming
- Emits tokens as they are generated, so answers appear progressively rather than all at once.
- Open weights
- The parameters are published, so the model can be inspected, fine-tuned and self-hosted under its licence.
- Caching
- Reuses already-processed prompt prefixes across requests, cutting input cost for repeated context.
Strengths & weaknesses
Strengths
- Tool-calling reliability under long loops
- Open weights
- Strong coding for an agent model
Weaknesses
- Prose style is utilitarian
- No vision in the base model
Pricing
Free| Rate | Price | Unit |
|---|---|---|
| Input | $0.00 | per 1M tokens |
| Output | — not charged | — |
Prompt caching is supported, so repeated prefixes bill below the listed input rate. Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.
Context window
200Ktokens
covers prompt and response together, so a long input leaves less room for the answer.
Supported features
| Feature | Support |
|---|---|
| Tool calling | Supported |
| JSON mode | Supported |
| Streaming | Supported |
| Vision input | Not supported |
| Audio input | Not supported |
| Long context | Not supported |
| Open weights | Supported |
| Reasoning | Not supported |
| Fine-tunable | Not supported |
| Batch | Not supported |
| Caching | Supported |
Example use cases
Autonomous agents
Tool-calling reliability under long loops
Computer-use workflows
Open weights
Tool-heavy backends
Strong coding for an agent model
Example prompt
Plan and execute: find the three cheapest 200K+ context models in this catalogue and produce a comparison table.Comparison
| Attribute | GLM-4.6 | Qwen3 Coder | Kimi K2 |
|---|---|---|---|
| Context window | 200K | 256K | 256K |
| Input price | $0.00 | $0.00 | $0.00 |
| Output price | — not charged | — not charged | — not charged |
| Tool calling | Supported | Supported | Supported |
| JSON mode | Supported | Supported | Supported |
| Vision input | Not supported | Not supported | Not supported |
| Open weights | Supported | Supported | Supported |
| Prompt caching | Supported | Not supported | Supported |
Best value in each row is highlighted. Illustrative placeholder data.
Provider information
View Zhipu AI →Zhipu AI · 智谱AI · Beijing, China · est. 2019
Born from Tsinghua University research, Zhipu's GLM family became the open ecosystem's agentic specialist: models tuned for tool use and multi-step work, released under permissive licences. Its free Air tier made capable open models accessible to every developer.
Recent releases
Documentation
Frequently asked questions
Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Prompt caching bills repeated prefixes below the listed input rate. Every figure on this page is an illustrative placeholder.