AI Briefing
KOSign in

Seahaven: Open-Source Framework for Building Stateful RL and Eval Environments

·2026.10.08 23:19

Key point

Seahaven provides isolated, stateful environments with SQLite backends and fixture-based resets for evaluating and training AI agents.

Details

Seahaven is a new open-source (MIT) framework designed to simplify the creation of realistic, stateful environments for evaluating and training AI agents. The tool addresses the difficulty of building reproducible eval environments by automating the infrastructure layer, allowing developers to focus solely on world-specific logic.

Core Capabilities

The framework manages the complexity of long-running agent tasks by providing:

  • Isolated Instances: Each run gets a private SQLite database copied from a fixture in milliseconds, allowing agents to modify state without affecting other runs.
  • State Diffs: Every row changed by the agent is logged, enabling grading based on actual world changes rather than transcript analysis.
  • Reproducibility: Runs are deterministic when using the same fixture, clock, and random seed.
  • Parallelism: Supports hundreds of instances per process for high-throughput evaluation.

Example Implementation: Stripe World

To demonstrate its capabilities, the author built a mock Stripe environment named Stripe World. This implementation includes 24 tables and 155 API operations, compatible with the official Stripe SDK and MCP server tools. This allows for realistic testing of financial agent behaviors in a sandboxed, resettable context.

Integration and Usage

Seahaven is compatible with OpenEnv, allowing integration with tools like Kiln auto-optimize and TRL's OpenEnv support. It supports publishing worlds to Hugging Face and serving environments via MCP or a web console. The framework requires Python 3.14+ and runs locally without external services.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.