CivBench: GLM-5.3 Leads Opus-5.5 in Civilization V, Qwen-3.8-27B Secures Science Victory
Key point
The updated CivBench evaluation shows GLM-5.3 ahead of Opus-5.5 in Civilization V, with both achieving Cultural victories, while Qwen-3.8-27B secured a Science victory.
Details
The updated CivBench framework evaluates large language models on their ability to play full games of Civilization V, a turn-based strategy game that tests long-horizon decision-making where consequences may not appear for 50 or 100+ turns. In the latest controlled runs, GLM-5.3 is reported as ahead of Opus-5.5, while Qwen-3.8-27B demonstrated strong capabilities.
Benchmark Methodology
The evaluation uses a controlled environment where models rotate through the same three fixed starts. Each game features eight civilizations: two controlled by the tested LLM strategist and six by the standard Vox Populi AI. The LLM sets high-level strategy, while the game's existing AI handles low-level execution.
Key Results and Accessibility
- GLM-5.3: Achieved a Cultural victory as China.
- Opus-5.5: Achieved a Cultural victory as Morocco.
- Qwen-3.8-27B: Secured a Science victory as China, demonstrating strong capability for its parameter count.
The project is open-source via Vox Deorum, allowing users to play against LLM-powered civilizations or watch AI-vs-AI matches. It supports local OpenAI-compatible servers and existing subscriptions for Claude or Codex, with estimated API costs of about $0.5 per player per game.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.