AI Briefing
KO

Qwen3.7-Max: agent frontier

·2026.05.21 09:33

Key point

Qwen3.7-Max showed top-tier performance on coding, reasoning, and agent benchmarks.

Details

Alibaba's Qwen3.7-Max has been released as an agent-focused proprietary model. It targets coding, office automation, and long-horizon autonomous execution, and was directly compared against competing models across multiple benchmarks.

Key performance results are as follows.

  • Scored 69.7 on Terminal Bench 2.0-Terminus, surpassing DS-V4-Pro Max.
  • Showed strong reasoning performance with 92.4 on GPQA Diamond.
  • Recorded high scores in general-purpose agent and office automation tasks as well, including 60.8 on MCP-Mark, 76.4 on MCP-Atlas, and 87.0 on SpreadSheetBench-v1.
  • Multilingual evaluation results were also presented, including 85.8 on WMT24++ and 89.2 on MAXIFE.

The training and evaluation approach focuses on generalized agent capability rather than simple score competition.

  • Training instances were separated into Task / Harness / Verifier, so the model was trained to work well across different combinations of harnesses and verifiers.
  • The benchmarks are explained as consisting of out-of-domain environments not included in training.
  • The approach claims to induce generalizable problem-solving strategies rather than shortcuts specific to a particular harness.

The most notable case is long-horizon autonomous optimization in a real-world environment.

  • The model autonomously optimized SGLang's Extend Attention kernel for 35 hours on a T-Head ZW-M890 PPU environment it had never seen during training.
  • During this process, it performed 1,158 tool calls and 432 evaluations, ultimately achieving a geometric mean 10.0x speedup compared to Triton, as stated.
  • The model continuously carried out kernel redesign, bottleneck analysis, correctness fixes, and memory optimization.

The announcement also emphasized applicability to reward hacking detection and long-horizon planning tasks.

  • During RL monitoring, it flagged 1,618 hacking cases and added 13 new heuristic rules, as explained.
  • On YC-Bench, it reportedly achieved $2.08 million in total revenue in a one-year startup simulation.

On the deployment side, an API will soon be available through Alibaba Cloud Model Studio, and integration methods with Claude Code, OpenClaw, and Qwen Code were also provided. This is an announcement positioning the model as a general-purpose AI agent model covering agents, coding assistants, office automation, and even robot navigation.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.