GLM 5.2 Ranks 3rd on ProgramBench
Key point
GLM 5.2 officially ranked 3rd on the ProgramBench benchmark, which tests reconstructing source code from binary executables.
Details
ProgramBench is an ultra long horizon benchmark where AI agents must reconstruct original source code using only a binary executable and a README.
GLM 5.2 ranked 3rd on the official leaderboard. This is considered a notable achievement given that many open-source and closed-source models have not yet been added to the current leaderboard.
Notably, on the cmatrix instance, it passed 99.8% of tests, coming close to a complete solve. The only failure was that it did not output the "Computer locked." string to the screen when running in -L lock mode.
The full agent trajectories will be published on programbench.com soon.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.