Skip to content

GLM-4.6

zhipu/glm-4.6
Open weightsAgents

The open agentic specialist, engineered for tool loops, computer use and multi-step autonomy.

Model overview

Context window

200K

tokens

Input price

quoted on request

Output price

quoted on request

Weekly volume

tokens / week

Open weights

Yes

self-hostable

Capabilities

5

of 11 tags

Description

GLM-4.6 is an agents model from Zhipu AI. Webparam does not stock this model yet and can source it on request — tell us what you need it for and we will come back with availability and a price. Zhipu AI’s stated focus is open agentic GLM family. Its mean it can also be self-hosted under its licence.

Capabilities

Tool calling
Calls functions you define and returns their arguments as structured data, the basis of agents.
JSON mode
Constrains the response to valid JSON matching a schema you supply.
Streaming
Emits tokens as they are generated, so answers appear progressively rather than all at once.
Open weights
The parameters are published, so the model can be inspected, fine-tuned and self-hosted under its licence.
Caching
Reuses already-processed prompt prefixes across requests, cutting input cost for repeated context.

Strengths & weaknesses

Strengths

  • Tool-calling reliability under long loops
  • Open weights
  • Strong coding for an agent model

Weaknesses

  • Prose style is utilitarian
  • No vision in the base model

Pricing

Free
Pricing for GLM-4.6, per 1M tokens
RatePriceUnit
Input$0.00per 1M tokens
Output not charged

Prompt caching is supported, so repeated prefixes bill below the listed input rate. Free at the point of use; fair-use rate limits apply to the free tier. Illustrative placeholder pricing.

Context window

200Ktokens

covers prompt and response together, so a long input leaves less room for the answer.

200K tokens against a peer range of 4K to 10M, median 200K. Compared across all models. This category has too few peers to scale against. Logarithmic scale.

Supported features

Feature support for GLM-4.6
FeatureSupport
Tool callingSupported
JSON modeSupported
StreamingSupported
Vision inputNot supported
Audio inputNot supported
Long contextNot supported
Open weightsSupported
ReasoningNot supported
Fine-tunableNot supported
BatchNot supported
CachingSupported

Example use cases

  • Autonomous agents

    Tool-calling reliability under long loops

  • Computer-use workflows

    Open weights

  • Tool-heavy backends

    Strong coding for an agent model

Example prompt

prompt
Plan and execute: find the three cheapest 200K+ context models in this catalogue and produce a comparison table.

Comparison

GLM-4.6 compared with Qwen3 Coder and Kimi K2
AttributeGLM-4.6Qwen3 CoderKimi K2
Context window200K256K256K
Input price$0.00$0.00$0.00
Output price not charged not charged not charged
Tool callingSupportedSupportedSupported
JSON modeSupportedSupportedSupported
Vision inputNot supportedNot supportedNot supported
Open weightsSupportedSupportedSupported
Prompt cachingSupportedNot supportedSupported

Best value in each row is highlighted. Illustrative placeholder data.

Provider information

View Zhipu AI
Zhipu AI3 models

Zhipu AI · 智谱AI · Beijing, China · est. 2019

Born from Tsinghua University research, Zhipu's GLM family became the open ecosystem's agentic specialist: models tuned for tool use and multi-step work, released under permissive licences. Its free Air tier made capable open models accessible to every developer.

Recent releases

    Documentation

    Frequently asked questions

    Input is billed at $0.00 per 1M tokens, and there is no separate output charge. Prompt caching bills repeated prefixes below the listed input rate. Every figure on this page is an illustrative placeholder.