[NAACL 2025 Best Paper Award] BiGGen Bench: An LLM Evaluation Framework Reflecting Precise Human Evaluation Criteria
Key point
LG AI Research developed BiGGen Bench, a new benchmark that reflects precise human evaluation criteria in LLMs, winning the NAACL 2025 Best Paper Award.
Details
Existing LLM benchmarks relied on abstract criteria such as preference or helpfulness, limiting their ability to precisely distinguish model performance. To solve this problem, LG AI Research's Super Intelligence Lab developed BiGGen Bench, which enables LLMs to leverage human experts' detailed judgment criteria in automatic evaluation.
BiGGen Bench built 77 detailed tasks, 775 instances, and specific Scoring Rubrics to diagnose 9 core capabilities of LLMs (Instruction Following, Grounding, Reasoning, etc.). This enables precise evaluation of logical validity or computational accuracy, going beyond simple preference measurement.
This research proved its academic value by being selected as the Best Paper Award, given to only one paper, at NAACL 2025, the most prestigious conference in the natural language processing field. This achievement is regarded as significant enough to be compared with past award-winning cases such as ELMo or BERT.
In actual performance verification, LG AI Research's EXAONE 3.5 model recorded an average score of 4.189 on the BiGGen Bench standard, demonstrating top-tier performance among the latest Non-Thinking models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.