JetBrains reports 16-point Python pass rate jump with Claude Fable 5
Key point
Claude Fable 5 achieved a 44.3% Python pass rate in JetBrains' private repository evaluations, surpassing Opus 4.8's 28.2%.
Details
JetBrains CTO Vladislav Tankov detailed how the company evaluates frontier models against its private repositories, including its monorepo, to ensure real-world performance matches benchmark scores. The company maintains leaderboards for quality, cost per task, and speed, noting that while Claude Fable 5 has a higher per-token cost, it often yields a lower cost per task for complex, long-running work.
Evaluation Results
In JetBrains' internal suite, Claude Fable 5 posted the best Python pass rate at 44.3%, a 16-point increase over Opus 4.8's 28.2%. In head-to-head comparisons, Claude Fable 5 solved 18 Python tasks that Opus 4.8 missed while losing only 2. The model also demonstrated greater efficiency, requiring approximately 22% fewer steps than Opus 4.8 to reach solutions and avoiding unnecessary attempts to pull in outside resources during Java tasks.
Deployment and Safety
JetBrains uses Claude Fable 5 for tasks requiring high-level reasoning, such as implementing complex components like rich text editors or rewriting applications across different runtimes. The company also employs the model for white-box security testing to identify vulnerabilities in its own products, preparing for a landscape where external actors may use similar models to probe for weaknesses. Tankov emphasized that JetBrains relies on Anthropic's red teaming for model safety while focusing its own efforts on infrastructure and harness security, accepting limited data retention for serious case investigations as a necessary trade-off for accessing frontier intelligence.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.