AI Briefing
KO

Agent Benchmark Testing Long-Horizon Terminal Task Performance (GitHub Repo)

·2026.07.14 09:00

Key point

The LHTB benchmark has been released to measure whether LLM agents can perform complex tasks spanning hundreds of steps in containerized terminal environments.

Details

Long-Horizon Terminal-Bench (LHTB) is a 46-task benchmark that measures whether LLM agents can sustain useful work over hundreds of steps in containerized terminal environments.

Unlike existing short-horizon coding benchmarks that simply produce a single output and terminate, LHTB places agents in a stateful environment. Evaluation is conducted rigorously through hidden, rebuild-from-artifact verifiers, rather than relying on the agent's self-reporting.

The main task scope is as follows:

  • Interactive games and puzzles
  • Multimodal analysis
  • Software and reverse engineering
  • Scientific computing and research reproduction
  • Security, performance, earth and energy systems
  • Professional APEX-style workflows

As of the July 2026 test results, Grok 4.5 recorded the highest average reward, but even the strongest model solved only about 28% of tasks under strict criteria. Notably, all models failed to solve the median task, showing that LHTB remains an unconquered territory.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.