AI Briefing
KO

[ACL 2024] Latest Trends and Key Insights in LLM Research - LG AI Research BLOG

·2026.07.16 09:00

Key point

Through the ACL 2024 tutorial session, we explore the latest research trends, including how to evaluate LLM-generated text and its vulnerabilities.

1 / 2

Details

With the advancement of generative AI, research to address the limitations of LLM(Large Language Model) and develop trustworthy models is actively underway. In particular, discussions on how to evaluate text generated by LLMs are at the core.

Existing N-gram based methods (BLEU, ROUGE) have limitations in that they are vulnerable to text variation and fail to capture semantic information. To address this, embedding-based methods such as BERTScore have emerged, but limitations in grasping context still remain.

Recently, the RLHF method, which trains a reward model through human feedback, and the LLM-as-a-judge method, in which the LLM itself becomes the evaluating agent, have drawn attention. MT-Bench is a representative LLM-judge methodology, offering high scalability and explainability, but it can suffer from model bias issues.

LG AI Research introduced two major studies as achievements in this field.

  • Prometheus 2: An evaluation-specific model that simultaneously utilizes Pairwise Ranking and Direct Assessment, showing high correlation with GPT-4 and human preferences.
  • Multi-Objective Reward Modeling: Reward model research designed to well reflect human preferences across various objectives.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.