Z.ai has published GLM-5 as an open-weight model aimed at long-horizon agentic and coding tasks. This revision treats the launch as a model-news story, not as an endorsement or an independent benchmark result from AI Model Watch.
The short version: GLM-5 is worth watching because Z.ai reports a much larger Mixture-of-Experts model than GLM-4.5, official API pricing that is lower than many frontier-model APIs, and benchmark claims that put it in the conversation for coding and agentic work. Those claims still need careful labels, because several numbers come from provider material or benchmark reports rather than our own testing.
What Z.ai says GLM-5 is
In its developer documentation, Z.ai says GLM-5 scales from GLM-4.5?s 355B total parameters and 32B active parameters to 744B total parameters and 40B active parameters. The same documentation says pretraining data increased from 23T to 28.5T tokens and that GLM-5 integrates DeepSeek Sparse Attention to reduce deployment cost while preserving long-context capacity. Source: Z.ai GLM-5 documentation.
The Hugging Face model card repeats the same high-level model facts and links to the API platform, paper, and GitHub repository. It frames GLM-5 around complex systems engineering and long-horizon agentic tasks. Source: Hugging Face GLM-5 model card.
Benchmark claims should be read as reported results
The model card reports several benchmark numbers, including 77.8 on SWE-bench Verified and a table comparing GLM-5 with GLM-4.7, DeepSeek-V3.2, Kimi K2.5, Claude Opus 4.5, Gemini 3 Pro, and GPT-5.2. These are useful reference points, but they should be described as reported benchmark results unless independently reproduced. Source: Hugging Face GLM-5 benchmark table.
The GLM-5 paper also reports a score of 50 on the Artificial Analysis Intelligence Index v4.0 and compares GLM-5 against other models across agentic, reasoning, and coding benchmarks. For publication quality, that should be phrased as ?the paper reports? rather than as a claim verified by AI Model Watch. Source: GLM-5 paper on arXiv.
Pricing: official API numbers
Z.ai?s pricing page lists GLM-5 at $1.00 per 1M input tokens, $0.20 per 1M cached input tokens, and $3.20 per 1M output tokens. These are official list prices from Z.ai?s documentation, but practical cost can still vary depending on provider, caching, routing, and usage pattern. Source: Z.ai pricing documentation.
Because pricing varies by provider and changes over time, this draft avoids broad claims like ?15 times cheaper? unless the comparison model, date, and source are explicitly included. A safer framing is that GLM-5?s official API price is aggressive for a large open-weight model, but readers should verify current pricing before making a buying or deployment decision.
License and access
The Hugging Face model page lists GLM-5 with an MIT license. That is useful for builders evaluating local or self-hosted workflows, but licensing and deployment feasibility are separate questions. A model can be open-weight while still requiring serious hardware, careful inference planning, and operational review. Source: Hugging Face GLM-5 model card.
What remains unverified by AI Model Watch
AI Model Watch has not independently run GLM-5 benchmarks, verified the training stack, validated data-privacy claims, or tested real-world API latency. Claims about healthcare, finance, education, or privacy-sensitive deployment should be treated as possible use cases, not proven outcomes, unless supported by a specific evaluation or deployment case study.
Hardware-independence claims also need careful sourcing. They may be important in the broader AI market story, but this article should not present them as proven from the current draft unless an official source or reliable reporting is added and clearly cited.
AI Model Watch take
GLM-5 is a strong candidate for close tracking because it combines large open-weight model scale, a coding-and-agentic positioning, official API pricing, and reported benchmark results that invite comparison with proprietary systems. The clean way to cover it is not hype; it is source-labeled analysis: what Z.ai reports, what third-party benchmarks report, what pricing pages currently show, and what still needs independent testing.
Source notes
- Z.ai GLM-5 documentation ? provider documentation for model scale, training-token claim, and DeepSeek Sparse Attention framing.
- Hugging Face GLM-5 model card ? model card, benchmark table, license, API/paper/repository links.
- Z.ai pricing documentation ? official GLM-5 API prices per 1M tokens.
- GLM-5 paper on arXiv ? paper/source for benchmark framing and Artificial Analysis Intelligence Index claim.