The End of the Cloud
·2026.04.15 04:05
Key point
Running MiniMax M2.7 AWQ on 2 GX10s to complete a local agentic coding stack.
Details
Connected 2 Asus Ascent GX10 units to expand a local inference setup, and directly tested several models for agentic coding.
- Initially felt that 1 unit with 128GB wasn't enough, so added a second GX10.
- Models tested were Qwen 3.5 122B-A10B, Qwen3-Coder-Next, M2.5-REAP, Qwen 3.5 397B-A17B, MiniMax M2.5 AWQ, and MiniMax M2.7 AWQ.
- Qwen 3.5 397B-A17B wasn't bad, but fell short as a reliable agentic teammate, and tended to declare results finished too quickly.
- MiniMax M2.5 AWQ is slow and has no vision, but was a very stable workhorse for agentic tasks, and MiniMax M2.7 was rated as an even better fit.
- In particular, when hooked up to a verification loop with tests or playwright-cli, it produced good results in planning, issue understanding, feature development, and bug fixing.
- However, one shouldn't expect the thoroughness of GPT-5.4 level or the strong watchdog tendencies of Opus 4.6, and that gap needs to be accounted for.
In conclusion, the author judged that the MiniMax M2.7 AWQ + 2 GX10 combo is enough to build a sufficiently practical agentic coding environment locally, and says they're now relying less on cloud providers.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.