Jeff v1.3 Release: Jeff-Code Accelerates Qwen 3.8-27B Coding Tasks by 47% While Maintaining Pass Rates
Key point
The new Jeff-Code agent uses a small decision model to pre-fetch information and manage thinking levels, reducing average task time to 0.68x of the baseline with no significant drop in accuracy.
Details
Jeff v1.3 introduces Jeff-Code, a coding agent fork of Pi that integrates a small decision model (Jeff) to optimize the performance of Qwen 3.8-27B. The core innovation is a System 1 approach where Jeff handles information-gathering steps (like reading files or listing directories) ahead of Qwen and dynamically decides whether Qwen needs to engage in deep thinking for each turn.
Performance Gains
In paired evaluations across 1,242 tasks, Jeff-Code reduced the average time per task to 0.68x of the baseline (Qwen 3.8-27B alone with full thinking), representing a 47% speedup (32% less time). Crucially, this speedup came with no statistically significant change in pass rates:
- Pooled Pass Rate: 62.4% (Jeff-Code) vs 62.8% (Baseline), difference -0.2 points (95% CI: -2.6 to +2.1).
- SWE-bench Verified: Time ratio 0.63x, pass rate +0.4 points.
- Terminal-Bench Pro: Time ratio 0.64x, pass rate +1.5 points.
- SWE-rebench: Time ratio 0.66x, pass rate -0.8 points.
Turning off thinking entirely made tasks faster but degraded quality by 7.6 points on average, demonstrating that Jeff's calibrated decision-making is essential for maintaining accuracy while saving time.
How It Works
Jeff operates via small LoRA adapters on a fixed base model, making two types of decisions:
- Pre-fetching: Jeff executes information-gathering tools (read, list, search) before Qwen's turn, providing results immediately. Writing and editing files remain with Qwen.
- Thinking Control: Jeff assigns a calibrated probability to whether deep thinking is needed. If confidence is below a threshold (0.6), thinking is skipped or reduced, saving significant latency.
Jeff v1.3 Base Model Updates
The underlying Jeff base model (v1.3) has been restructured to prioritize prefix caching over zero-shot performance. By placing fixed instructions first and live data last, repeated decisions are faster. While zero-shot accuracy on long option lists dropped significantly (e.g., legal-clause classification fell from 66.0% to 7.4%), accuracy is restored via adapters. With the appropriate adapter, v1.3 performance is within ±0.7 points of v1.2.
Adapters and Availability
The release includes 15 adapters (9 retrained, 6 new), including two for Jeff-Code and four for real-world tasks (trading-desk, AML, sanctions, SOC). GGUF quantizations (Q8_0 and Q4_K_M) are available for llama.cpp, with Q8_0 being effectively lossless. The project is open-source under MIT license (for the Pi fork) and hosted on GitHub and Hugging Face.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.