Local coding models reach practical-use level
Key point
On Terminal-Bench 2.0, 27B-32B open-weight coding models recorded up to 38.2%.
Details
Evaluating 27B-32B open-weight coding models on Terminal-Bench 2.0 (89 tasks), Qwen 3.6-27B scored the highest at 38.2% (34/89) under the default per-task timeout.
- The comparison baseline is the default setup, the same as the public leaderboard.
- The author stated that these figures were compared against the verified leaderboard.
- It was noted that MOE models still have more than single-digit room for improvement in inference speed.
The verified SOTA on the public leaderboard is around 80%, and the post interprets the performance of the latest runnable offline models as comparable to hosted frontier models from late 2025.
- Terminus 2 + Claude Opus 4.1: 38.0%
- Terminus 2 + GPT-5.1-Codex: 36.9%
- Claude Code + Sonnet 4.5: 40.1%
- Codex CLI + GPT-5-Codex: 44.3%
The post's conclusion is that coding models runnable in offline/on-premise/regulated environments have, for the first time, come close to the boundary of being practically deployable. The authors summarized this as roughly a 6-8 month performance gap.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.