AI Briefing
KO

Z.ai Releases GLM-5.3-Flash: Approaching Claude Opus 4.8 with 18B Active Parameters

·2026.08.27 09:00

Key point

Z.ai released GLM-5.3-Flash with 18B active parameters, achieving performance close to Claude Opus 4.8 at one-tenth the cost.

1 / 4

Details

Z.ai released GLM-5.3-Flash, the first native multimodal model in the GLM-5 series. With only 18B of the total 320B parameters active, it demonstrates superior performance over the previous GLM-5.2 in benchmarks and real-world workloads while costing one-tenth as much. In coding and agentic benchmarks, it recorded results approaching Claude Opus 4.8.

Efficient Architecture

GLM-5.3-Flash introduces a hybrid architecture combining sparse and linear attention to significantly reduce long-context serving costs. It also applies Manifold-Constrained Hyper-Connections (mHC) to improve scaling efficiency. Compared to the GLM-4.5 series, the total parameter count is similar (320B vs 355B), but active parameters (18B vs 32B) and layer count (45 vs 92) are nearly halved, optimizing for ultra-low-cost inference.

Performance and Cost Competitiveness

It scored 57 points on the Artificial Analysis Intelligence Index v4.1.1, providing an intelligence level that previously cost about 10 times more for $0.045 per token (with discount applied). It consistently outperformed GLM-5.2 across six coding and agentic benchmarks, showing significant gaps particularly in DeepSWE v1.1 (63.4 points vs 46.2 points) and AutomationBench (48.8 points vs 26.2 points).

Pre-release Testing and Infrastructure

Before release, it was tested under the anonymous name ox-alpha on OpenCode and OpenRouter, with all traffic processed on Chinese AI chips. To support a 1M token context length, it introduced IndexPool technology to reduce indexer latency and memory overhead.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.