AI Briefing
KO

Mobile Agent Benchmark Released

·2026.09.07 15:39

Key point

A new mobile agent benchmark named 'MobileWorld' has been released to evaluate capabilities in user interaction and MCP tool utilization.

Details

A new benchmark 'MobileWorld' has been released to evaluate the performance of LLM agents that autonomously perform GUI tasks on mobile devices. To address the limitations of existing evaluation methods, this benchmark introduces two new evaluation axes: User Interaction and MCP (Model Context Protocol) tool utilization.

Benchmark Composition and Features

  • Evaluation Scope: Includes a total of 201 tasks targeting more than 20 commonly used apps, such as communication, messaging, and productivity. System apps account for less than 5%, configured to closely resemble real user environments, though some apps use open-source alternative versions, which presents certain limitations.
  • Architecture: Uses a Planner-Executor structure consisting of a Planner based on a VLM (Vision-Language Model) that receives only screenshots as input, and a Grounding Model that converts natural language instructions into coordinates (x, y).
  • New Evaluation Axes:
    • User Interaction Tasks: Evaluates the agent's ability to ask the user questions and receive answers to complete tasks in situations with insufficient information. Here, the 'user' is simulated by the GPT-4 model.
    • MCP Tasks: Evaluates the ability to perform tasks more efficiently than GUI clicks by utilizing MCP tools capable of external data collection and complex operations, such as GitHub and arXiv.

Performance Results

The combination that recorded the highest performance was Gemini-3-Pro combined with UI-Inst-7B, with an average accuracy of approximately 52%. In contrast, end-to-end (E2E) GUI-only models showed significantly lower performance on the two new evaluation axes. This reveals the limitations of existing GUI agents in complex data collection and interactive tasks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.