AI Briefing
KO

Generalization Performance Results from Training on the APEX-Agents Development Set

·2026.04.02 09:00

Key point

**AC-Small**, trained on the APEX-Agents development set, demonstrated remarkable generalization performance improvements on major benchmarks.

1 / 2

Details

The AC-Small model, created by post-training GLM-4.7 on the APEX-Agents development set, was verified to determine whether it achieved genuine generalization of capabilities beyond simple benchmark score improvements.

The verification results showed substantial performance improvements across held-out industry benchmarks:

  • APEX: +5.7 point improvement
  • Toolathalon: +8.0 point improvement
  • GDPVal: +7.7%p improvement

On the GDPVal benchmark, AC-Small rose from 7th place to 5th place, surpassing Claude Opus 4.5. This indicates a substantive improvement in the ability to perform economically valuable professional work.

On Toolathalon, tool-use capability improved significantly, recording 34.6%, surpassing Claude-4-Sonnet, Kimi-K2.5, Grok-4, and others, with a 4-rank rise.

On the APEX benchmark, professional reasoning capability improved, and notably, despite being trained only on legal, consulting, and financial data, the model achieved a surprising 4th place ranking in the Medicine field. This is because the work process itself became more sophisticated, such as the model accurately describing details without omission when writing diagnostic reports.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.