GPT-5.6 Luna finds 75% of verified bugs at 3.6% of GPT-6 Astra's cost, but with lower precision
Key point
GPT-5.6 Luna demonstrated high cost efficiency by finding 75% of verified bugs at only 3.6% of the cost of GPT-6 Astra, but showed inferior performance in precision and security vulnerability detection.
Details
According to a benchmark published by Aditya Jha of Entelligence, GPT-5.6 Luna used only 3.6% of the cost of GPT-6 Astra to find 75% of verified bugs. Across 50 PRs, Luna's total cost was $0.20 and Astra's was $5.66, with a per-PR cost of $0.0041 for Luna and $0.113 for Astra, a difference of approximately 28x. The cost per verified bug was $0.0030 for Luna and $0.061 for Astra.
In terms of precision, Luna scored 74% and Astra 96%, indicating a higher error rate for Luna. Specifically, in security vulnerability detection, Luna found only 9 out of 24, while Astra found 19, showing a significant gap. In the Keycloak repository, Luna's precision dropped to 50%. In reproducibility tests, Luna showed a reproducibility rate of 47% and Astra 67%. Overall, Luna was assessed as suitable for general logic bugs but unsuitable for code reviews involving security and complex permission models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.