AI Briefing
KO

Open RL Recipe TMax for Terminal Agents Released

·2026.06.23 00:38

Key point

A new RL recipe and a large-scale dataset, TMax-15k, designed to maximize terminal agent performance have been released.

Details

TMax, a new reinforcement learning (RL) recipe designed to push terminal agent performance up to frontier-model levels, has been released.

The key highlights are as follows:

  • TMax-15k Dataset: A dataset of 14,600 RL environments built through a compositional pipeline that allows precise control over difficulty and diversity. This is more than 2.5x larger than the largest existing open terminal dataset.
  • RL Recipe: A simple Outcome-only RL approach is proposed, combining GRPO with stability-improving techniques.

Performance Metrics:

  • TMax-9B: Scored 27.2% on Terminal Bench 2.0, achieving top-tier performance among models under 10B. This is close to Claude Haiku 4.5's score of 29.8%.
  • TMax-27B: Performance improved to 42.7%, showing results close to those of massive models like Kimi K2.5 (43.2%).

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.