AI Briefing
KO

GPT-6 Astra Achieves 95% Success Rate and 2.3x Cost Reduction vs. Claude Fable in Robot Manipulation Tasks

·2026.09.04 09:00

Key point

GPT-6 Astra outperformed Claude Fable 5.1 in robot block insertion tasks with a 95% success rate and 2.3x lower cost.

Details

OpenAI's GPT-6 Astra demonstrated overwhelming performance and efficiency compared to Anthropic's Claude Fable 5.1 in robot manipulation benchmarks. In the Bowl task (inserting blocks into a bowl), Astra recorded a 95% (19/20) success rate, while Fable 5.1 achieved 40% (8/20) and Fable 5 only 5% (1/20).

Performance and Cost Efficiency

Astra held a significant advantage not only in task completion rates but also in resource usage. For the Bowl task, Astra's average time was 2.5 minutes, token usage was 2.1k, and cost was $0.94. This represents a 2.4x higher completion rate and 2.3x lower cost compared to Fable 5.1's 6.8 minutes, 12.9k tokens, and $2.12 cost.

Limitations in Complex Tasks

In contrast, both models showed low success rates in the Puzzle task (fitting puzzle pieces into slots). Both Astra and Fable 5.1 recorded a 10% (2/20) success rate, while Fable 5 was at 0%. However, Astra still consumed fewer tokens (2.7k vs 12.9k) and lower costs ($1.36 vs $2.18) compared to Fable 5.1.

Experimental Environment and Limitations

The experiments were conducted using the Inspect Robots 0.58.0 harness and YAM arms (6-DoF, parallel grippers). Each model performed 20 attempts per task, with policies set to 'medium thinking' and a budget of 20 LLM calls. Limitations noted include the time lag between Astra and Fable executions, the asymmetry of the robot rig used for the Bowl task comparison, and the potential for unconscious bias from evaluators.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.