AI Briefing
KOSign in

Fig Analysis Reveals Jagged Performance in Frontier Models Across Agentic Tasks

·2026.09.30 09:00

Key point

A new analysis by Fig demonstrates that frontier models like Astra and Opus 5.5 exhibit jagged performance across web and physical tasks, with no single model consistently outperforming others.

1 / 4

Details

Fig evaluated frontier models on web automation (VisualWebArena) and physical domains (Bench2Drive, Assembly101, IndEgo, VLABench). For web tasks, seven models including Fable 5, GPT-5.6, Opus 4.7, Opus 5, Opus 5.5, Kimi K3, and GPT-6 Astra were tested. For physical tasks, three models (Gemini 3.5-Flash, GPT-6 Astra, Opus 5.5) were evaluated. The study found no clear performance leader; in every model pair, the lower-scoring model solved tasks the higher-scoring one missed. Performance varied significantly by website, with Astra scoring highest on Reddit and Fable 5 on Classifieds. Benchmark difficulty labels showed weak correlation with model success. Model updates reshaped performance profiles without necessarily moving the average score; for instance, the update from Opus 5 to Opus 5.5 raised the mean score by 5.1 percentage points but still lost 11 tasks that Opus 5 had solved. The researchers released the RIDGE dataset containing item-level results.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.