AI Briefing
KOSign in

LMGame Team Demonstrates Hybrid AI Swarm Architecture Outperforms Lone Models in Detective Game

·2026.10.11 15:35

Key point

A hybrid system combining GPT6-Astra planning with a Jev agent swarm achieved 96% success on standard cases, significantly outperforming the lone model's 57%.

Details

The LMGame team developed a procedural detective game set in Victorian London to evaluate the efficacy of AI agent swarms versus single large models. The benchmark involves solving cases with 100 addresses, 10 suspects, and a single lead per address, requiring the identification of the culprit, motive, and hideout within 12 hours of simulated case time.

Performance Metrics

The study compared three architectures using GPT6-Astra as the reasoning engine and Jev agents for search tasks:

  • Holmes (GPT6-Astra) + Jev Swarm: 96/100 success on standard cases and 93/100 on frame-up cases (where witnesses lie).
  • Rule-based Holmes + Jev Swarm: 95/100 on standard cases but 0/100 on frame-up cases.
  • The Lone Genius (GPT6-Astra): 57/100 on standard cases and 59/100 on frame-up cases.
  • Jev Swarm (no leader): 38/100 on standard cases and 42/100 on frame-up cases.

Key Findings

  • Divide and Conquer: The swarm excels at evidence coverage (the divide step), while the leader model handles reasoning (the conquer step).
  • Swarm Vulnerability: Leaderless swarms converge too quickly on incorrect answers due to error propagation, despite being the fastest method (median 2.2 case hours).
  • Leader Necessity: A simple rule-based leader matches GPT6-Astra's performance on honest evidence but fails completely when faced with deception. A strong reasoning model is required to detect lies and identify the correct culprit in frame-up scenarios.
  • Lone Model Limitations: Single models struggle with spatial coverage, often missing critical clues located at addresses they do not visit within the time limit.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.