Benchmark Shows GPT-6 Luna's Confidence Scores Unreliable for Complex Decisions
Key point
A new benchmark reveals that GPT-6 Luna, the model powering OpenAI's upcoming Decisions API, exhibits severe overconfidence and accuracy drops on multi-step reasoning tasks.
Details
Researchers tested GPT-6 Luna, identified as the base for OpenAI's new Decisions API, on the Hard-Decisions benchmark. While Luna performs well on simple tasks, its accuracy falls to approximately 46% at five reasoning steps, compared to 85% for the competing model Jev. More critically, Luna's confidence scores are poorly calibrated: when the model reports 99% certainty, it is correct only about 68% of the time. The Area Under the Curve (AUROC) for Luna's ability to distinguish correct from incorrect answers is 0.68, significantly lower than Jev's 0.85+. The study suggests that unless OpenAI's specialized Decisions API version addresses these calibration issues, routing software decisions based on Luna's confidence may lead to frequent errors.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.