AI Briefing
KO

AI Startup Hark Launches First Product: Affordable and Fast Computer-Use Agent Hark Handoff (7-min read)

·2026.08.06 09:00

Key point

AI startup Hark has launched Handoff, a computer-use agent that automates web tasks.

Details

AI startup Hark has launched Hark Handoff, a computer-use agent (CUA) that navigates websites and performs tasks on behalf of users. Handoff autonomously handles tasks from start to finish, such as ordering food on DoorDash, booking flights on United and Delta, and sending messages to job seekers on LinkedIn. Public sign-ups have begun at Hark.com, and the initial software platform is scheduled to launch later this month.

According to Hark, Handoff scored 97.7 on the web agent benchmark Online-Mind2Web (OM2W).

  • Handoff: 97.7
  • OpenAI GPT 5.4: 92.8
  • Anthropic Claude Opus 4.8: 84.1
  • Google Gemini 2.5 Pro: 69

Model serving costs are $0.18 per million input tokens and $2.37 per million output tokens, lower than the $5 and $30 for GPT 5.5, as compared by Hark. The company stated that the model latency per turn is 0.8 seconds.

For each request, Handoff creates a dedicated virtual computer equipped with a browser, file system, and terminal. Users can connect their existing accounts to utilize saved addresses, payment methods, and usage history. Hark pointed out that while 75% of daily screen time is spent in browsers, fewer than 1 in 1,000 websites offer public APIs, making agent automation difficult.

However, there are limitations to the performance comparison. The OM2W comparison presented by Hark did not include the latest models such as OpenAI GPT-5.6, Anthropic Opus 5, DeepSeek V4, Kimi K3, and Qwen3.8-Max. Since the OM2W results for these models have not been released, Hark's claim of 'best ever' cannot be directly verified against the current strongest systems.

In Hark's own comparison, GPT 5.5 scored 72.3 on WebTailBench v2, surpassing Handoff's 68.6. Additionally, two of the WebTailBench and internal evaluations used Hark's own test environment and internal LLM evaluators, and the latency of competing models was also measured by Hark in its own environment using the slowest inference settings.

Price competitiveness is relatively clear. Since Anthropic's latest Opus 5 also costs around $5 per million input tokens and $25 per million output tokens, the cost savings offered by Handoff are likely to hold even when compared to the latest models. Hark stated that it trained the model by combining pre-supervised learning with asynchronous reinforcement learning using the GRPO algorithm.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.