GLM 5.1 running locally at 40tps
Key point
Ran GLM 5.1 at over 40tps on 4x RTX 6000 Pro.
Details
Ran the reap-ed nvfp4 version, patched through sglang, stably on a 4 x RTX 6000 Pro setup. Power was capped at 350W per card.
Prefill throughput was 2229.0 pp/s at 4096 tokens, and generation throughput was 42.03 tps at 512 tokens.
Performance gradually declined as context length increased, but it still maintained 863.5 pp/s / 35.87 tps even at 64K context.
- 0 context: 2229.0 pp/s, 42.03 tps
- 4K: 1943.6 pp/s, 41.41 tps
- 16K: 1558.9 pp/s, 39.72 tps
- 32K: 1234.2 pp/s, 38.19 tps
- 64K: 863.5 pp/s, 35.87 tps
The author rated the opencode experience as close to Sonnet + Claude Code, and noted that 100~200k sessions were also stable.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.