Step 5 Preview Released: 600B Parameter MoE Model Highlights Agentic Tasks and Financial Analysis Performance
Key point
Step 5 Preview delivers competitive performance in agentic tasks and financial analysis with a 600B parameter MoE architecture, achieving a score of 44 on the Artificial Analysis Intelligence Index.
Details
Step 5 Preview is a flagship model for agentic tasks, featuring a sparse MoE architecture with 600B total parameters (27B active per token). It supports a 1M token context window and vision input, achieving a score of 44 on the Artificial Analysis Intelligence Index and claiming lower task costs compared to models of similar intelligence levels.
Coding and Long-Running Performance
It recorded scores of 67.7 on DeepSWE v1.1, a software engineering benchmark, and 49.0 on StepCodeBench. In GPU kernel optimization tasks on NVIDIA H100, it achieved 508 TFLOPS within a 24-hour limit, surpassing Claude Opus 5 (493 TFLOPS). Additionally, in the Pokémon Red environment, it sustained over 3,000 turns and 6 million tokens of interaction without specialized optimization, progressing through approximately 1/3 of the main storyline.
Financial and Expert Knowledge Analysis
It scored 66.4 on the FrontierFinance benchmark, outperforming GLM-5.3 (64.1) and GPT-6 Astra (55.0), but falling short of Claude Opus 5 (69.7). It can perform 950 web fetches and collect 300,000 monthly records in a single agent action, executing complex domain-specific analyses such as generating workbooks that analyze 17 sheets when reviewing diesel surcharges.
Comparison with Competing Models
In overall reasoning and coding capabilities, there remains a significant gap compared to top-tier frontier models like GPT-6 and Claude Opus 5. It recorded lower scores than competing models in areas such as GPQA Diamond (93.5%) and DeepSWE v1.1 (67.7%).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.