Proactive Agent Research Environment: Evaluating Proactive Assistants via Active User Simulation
Key point
It proposes the Pare framework and benchmark that evaluate proactive AI agents through sophisticated simulation of user behavior.
Details
Proactive Agents, which anticipate user needs and autonomously carry out tasks, are core to digital assistants, but their development has been hampered by the absence of a realistic user simulation framework. Existing approaches treat apps as simple API call models, failing to properly capture the state changes and sequential interactions unique to digital environments.
To address this, we introduce the Pare (Proactive Agent Research Environment) framework. Pare is designed to model applications as Finite State Machines, enabling user simulators to perform state-based exploration and actions. This makes active simulation that closely resembles real users possible.
We also present Pare-Bench, built to verify agent performance. This benchmark covers a total of 143 diverse tasks spanning the following areas:
- Using communication and productivity apps
- Performing scheduling and lifestyle services
- Context Observation and Goal Inference
- Evaluating Intervention Timing and multi-app orchestration capabilities
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.