Opus 4.7 Rebounds
·2026.04.17 11:35
Key point
Opus 4.7 has clearly improved, while Kimi K2.6-Code-Preview remains a wait-and-see.
Details
We added coding evaluations, accumulating 145 results across the new and old version leaderboards.
The models tested in this pass include Kimi K2.6-Code-Preview, Opus 4.7, GLM 5.1, and Minimax M2.7, among others.
Key impressions are as follows.
- Opus 4.7 showed a substantial improvement, which is rare among recent upgrades.
- Kimi K2.6-Code-Preview doesn't yet feel significantly better, so judgment was reserved pending more use in other agentic environments.
- GLM 5.1 was quite good, and the author believes some open-weight models still don't reach Opus/GPT tier, contrary to exaggerated claims.
- In the top tier are Kimi K2.5 and GLM 5.1, which were assessed as possibly approaching Gemini/Sonnet level.
- In the mid tier, Minimax M2.7 and Qwen 3.6 Plus were named, and were considered still attractive in terms of price or local-run feasibility.
Another point of note is ForgeCode.
- It recorded the highest score with Minimax M2.7.
- However, its UX/DX differs greatly from OpenCode, making it hard to recommend.
- It comes in the form of a Zsh plugin, which may suit users who prefer that format, but at the time it was summarized as too buggy for immediate use, with breakage occurring across most other models/providers.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.