AI Briefing
KO

OpenAI GPT-6 Astra's 99.9% Score Driven by Harness Effects; ARC Prize Rejects AGI Declaration

·2026.09.06 21:23

Key point

The 99.9% score for GPT-6 Astra, presented by OpenAI as evidence of AGI, was achieved using a custom harness, while it scored only 62.7% under the standard harness, leading ARC Prize to reject the AGI declaration.

1 / 2

Details

OpenAI declared the dawn of the AGI era based on GPT-6 Astra's ARC-AGI-3 benchmark score of 99.9%, but data released by ARC Prize indicates this result stems from differences in the harness rather than the model's inherent capabilities. Under ARC Prize's standard harness environment, Astra's score was only 62.7%, whereas it achieved 99.9% only in the custom Adapter environment used by OpenAI.

Harness Scaffolding Has Greater Impact Than Reasoning Intensity

The score difference between the two harnesses arose from how the harness—the software surrounding the model—managed tools, memory, and context. Notably, when the reasoning setting was set to 'none' in the OpenAI adapter, Astra recorded 96.7%, which was 34 points higher than the score under the standard harness's maximum reasoning setting (62.7%). This suggests that scaffolding has a greater impact on performance than reasoning intensity.

Additionally, the OpenAI adapter environment was more cost-effective than the standard harness. For Astra, the cost under the standard harness (maximum reasoning) was $26,098, while the cost under the OpenAI adapter (high reasoning) was $18,817, achieving a higher score at a lower cost. The adapter includes opaque reasoning state maintenance and conversation compression features.

Rejection of AGI Declaration and Benchmark Reliability Controversy

ARC Prize officially rejected OpenAI's AGI declaration based on this score discrepancy. Co-founders Simon Willison and François Chollet pointed out the need to distinguish between the model's inherent capabilities and the performance of the assembled system. They noted that while the 'model' is what is sold, it was the 'assembled system' that achieved the 99.9% score.

The volatility of benchmark scores was also cited as a problem. ARC Prize changed scores after publication, and according to Fortune, Astra's hallucination score fluctuated drastically from 4.2% to 51%, with changes made without explicit documentation. Stanford researchers labeled this practice 'benchmaxxing', criticizing the behavior of rerunning tests with altered conditions until scores improved.

Gap Between Independent Verification and Model Capabilities

Independent testing by Artificial Analysis showed that Astra tied with Claude Opus 5 and Fable 5 (67 points) on the Coding Agent Index, but scored 61 on the Intelligence Index, 5 points lower than Fable 5.1. Pricing was also 2.5 times higher than Sol, at $10 per input token and $50 per output token.

Ultimately, the benchmarks revealed a gap between the model's actual capabilities and performance independently verifiable by third parties. OpenAI removed and then republished the post without disclosing the reason, explaining that evaluations contain noise of a few percentage points depending on checkpoints and execution environments.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.