AI Briefing
KO

Qwen Releases 'E-Commerce Bench' in Collaboration with Taobao to Evaluate Long-Term Operational Capabilities

·2026.09.03 11:00

Key point

The Qwen team, in collaboration with Taobao, released E-Commerce Bench, a long-term operational agent benchmark based on one year of real market data.

1 / 10

Details

The Qwen team, in collaboration with Taobao & Tmall Group, released E-Commerce Bench, which evaluates the long-term business operational capabilities of AI agents. While existing benchmarks focused on short-term tasks with clear endpoints, this benchmark measures continuous decision-making capabilities in real business environments that lack natural endpoints.

One-Year Simulation with Limited Resources

Agents operate multiple online stores for 365 days starting with initial capital of 100,000 yuan. Each day is limited to 600 minutes from 8:00 AM to 6:00 PM, and every tool call consumes time. For example, checking the balance takes 10 minutes, while opening a store requires 60 minutes. Under these time budget constraints, agents must perform category research, supplier negotiation, pricing, and inventory and cash flow management.

Deterministic Environment Based on Real Taobao Data

To ensure the benchmark's reliability, a Deterministic environment was built based on real data from the Taobao & Tmall platform.

  • Real Market Data: Includes 60 categories, 6,886 products, 576 suppliers, and 12 store types, with 10 market events and 8 promotions fixed on an annual schedule.
  • Three-Account Settlement System: Costs are deducted immediately from the bank account, but revenue must go through commission deduction after delivery, a 9-day platform escrow wait, and a manual withdrawal process to be cashed out.
  • Inventory and Shipping: Inventory incurs daily storage fees, and costs vary depending on shipping speed. Orders not shipped within 2 days are canceled.
  • Reputation and Returns: Return rates are determined by category benchmarks, supplier defects, pricing, shipping speed, etc. Reputation directly affects demand, and low reputation can reduce traffic by up to 15%.

Demand Model and Negotiation Kernel

To clearly distinguish performance differences between models, demand and supplier behaviors are calculated through named factors (base demand, price response, weekends, promotions, seasonality, etc.) rather than random sampling. This ensures the benchmark's reproducibility and discriminative power.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.