Replit Agent Harness Beats Sidekick Architecture by 11-16 Points; Pareto-Efficient Against Astra
Key point
Replit Agent's new harness design beats the sidekick architecture by 11 points on DeepSWE and 16 on Terminal-Bench, achieving Pareto-efficiency against Astra baselines.
Details
Replit Agent now allows the core model to decide effort levels, hand-offs, and specialist delegation, replacing static routing. On DeepSWE v1.1, Replit Agent scores 72% at $2.11 per task, beating the sidekick architecture (61% at $1.34) by 11 points. On Terminal-Bench 4.0, it reaches 49% at $2.53 per task, surpassing the sidekick architecture (33% at $1.84) by 16 points. Compared to Astra baselines, Replit Agent is Pareto-efficient: Astra's low-effort setting scores lower (67% on DeepSWE, 42% on Terminal-Bench) but costs less, while Astra's high-effort setting scores higher (74% on DeepSWE, 60% on Terminal-Bench) but costs more than twice as much. No Astra baseline wins on both cost and score simultaneously.
Composable Primitives for Delegation
The harness provides four primitives for the core loop:
- Domain-aware subagents: Specialists for exploration, testing, review, and design run on models strongest in those areas.
- Subagent tiers and effort: The core loop selects small, standard, or large tiers and adjusts effort dynamically.
- Reusable subagents: The model can return to previously briefed subagents, leveraging longer cache lifetimes on newer models.
- Dynamic effort tuning: An escalation system adjusts effort mid-turn based on task difficulty, preserving cache on models like GPT-6 Astra and Fable 5.1.
Model-Specific Delegation Behaviors
Production data shows distinct delegation patterns. GPT-6 Astra routinely delegates to general workers (20% of turns) and returns to existing subagents (42% of dispatches). In contrast, Fable 5 and Fable 5.1 rarely hand work to general workers (0.9% and 2.3% respectively), preferring to keep implementation tasks themselves.
The Bitter Lesson of Harness Design
Replit argues that composable harnesses allow models to scale with their capabilities, aligning with Sutton’s bitter lesson that general methods eventually outperform human-designed heuristics. By letting the model discover execution strategies, Replit Agent achieves better outcomes than rigid architectures.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.