The Rebellion of Local 35B
Key point
An experiment tackling SWE-bench Verified by running Qwen3.6-35B A3B on 2x Tesla P40.
Details
Qwen3.6 35B A3B was run with 4-bit quantization on two Tesla P40 cards to test the limits of a local coding agent.
The target was SWE-bench Verified, a benchmark that requires fixing bugs in real Python repositories to pass hidden tests.
The original text explains the starting point of the experiment: to find out why this benchmark is expensive, and whether a local model can still produce practical code patches.
As a point of comparison, it also mentions publicly disclosed figures from top-tier models, presenting context against existing performance tables such as GPT-5 Mini 56.2%.
The core point is to verify how much agentic coding performance can be pushed in a local environment with specific hardware and configuration, without relying on massive cloud models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.