AI Briefing
KO

[NAACL 2025 Best Paper Award] BiGGen Bench: A Principled Benchmark for Fine-Grained Evaluation of Language Models - LG AI Research Blog

·2026.07.16 09:00

Key point

LG AI Research developed BiGGen Bench, a new benchmark for precisely diagnosing the capabilities of language models, and won the NAACL 2025 Best Paper Award for it.

Details

Existing language model benchmarks relied on abstract criteria such as preference or usefulness, which limited their ability to finely distinguish model capabilities. To address this, the Super Intelligence Lab at LG AI Research collaborated with Professor Minjoon Seo's research team at KAIST and global universities to develop BiGGen Bench.

BiGGen Bench defines 9 core capabilities and 77 detailed task types for language models, and designed a total of 775 prompts along with corresponding evaluation rubrics. This research was recognized for its value by winning the Best Paper Award—the only one selected among over 2,000 papers—at NAACL 2025, a prestigious conference in the natural language processing field.

This benchmark proposes the following hybrid evaluation framework:

  • Human: Ensures reliability by defining detailed evaluation rubrics for each domain and task type
  • LLM: Ensures scalability by performing evaluation based on the defined rubrics

Additionally, when EXAONE 3.5, LG AI Research's LLM, was evaluated using this benchmark, it demonstrated outstanding performance by recording an average score of 4.189, the highest level among recent non-reasoning models, excluding reasoning models.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.