dig.bench (Website)
Key point
dig.bench is a benchmark of 70 text-based games measuring AI agents' scientific discovery capabilities.
Details
dig.bench is a benchmark platform that evaluates AI agents' ability to discover hidden game rules through experimentation. It consists of 70 text-based games, excluding visual elements to focus on language models' pure discovery capabilities.
All games are designed to be solvable by humans on the first attempt and are divided into 7 difficulty tiers. Currently, 21 games are released, and humans compete with state-of-the-art models using the same interface and step limits.
Evaluation is conducted via two methods: the Basic Harness and the Agentic Harness. The Basic Harness handles the next action and game state, while the Agentic Harness adds tools like a file system and context management to support more complex agent tasks.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.