AI Briefing
KO

AI Safety Benchmark DystopiaBench Released

·2026.05.18 22:03

Key point

DystopiaBench has been released, measuring how well AI models detect cleverly hidden dangerous requests.

Details

DystopiaBench is a benchmark that evaluates LLM safety through 36 step-by-step scenarios based on 6 dystopia types (Petrov, Orwell, Huxley, Basaglia, LaGuardia, Baudrillard).

This benchmark measures whether models can go beyond simply refusing obviously dangerous requests, and can recognize and refuse dangerous requests that are cleverly hidden through dual-use or normalization techniques (e.g., building a social credit system).

Key updates and features are as follows:

  • 42 models tested: Includes both open-source and closed models.
  • LLM-as-a-judge: Uses 3 LLMs as judges to calculate scores.
  • Improved precision: Uses the average of 3 runs as the final score.
  • Expanded modules: Added 4 new modules and scenarios in addition to the existing Petrov and Orwell.

This project has been released as open source, allowing anyone to contribute to or use it via GitHub.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.