AI Agent Benchmark TextQuests Released
Key point
TextQuests, a text-based game benchmark for evaluating LLMs' reasoning and exploration abilities as agents, has been announced.
Details
As static knowledge evaluation of LLMs becomes saturated, measuring performance as autonomous agents in dynamic, interactive environments has become important. TextQuests is a new benchmark that validates agent capabilities using 25 classic Infocom text-based games.
This benchmark requires agents to demonstrate the following two core capabilities:
- Long-Context Reasoning: The ability to establish and execute multi-step plans based on extensive action and observation history.
- Learning through Exploration: The ability to learn and improve from experience through trial and error.
Evaluation is conducted in two ways, with and without hints, using Game Progress and a Harm metric that measures the agent's ethical behavior.
Experimental results showed that when context exceeds 100K tokens, models tend to hallucinate about previous interactions or repeat past actions instead of forming new plans. This limitation was particularly pronounced in environments requiring spatial reasoning.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.