AI Briefing
KO

[NeurIPS 2023] RLHF, RLAIF Research Trends and Key Papers

·2026.07.16 09:00

Key point

This introduces the latest research trends and key papers on RLHF and RLAIF presented at NeurIPS 2023.

1 / 2

Details

Reinforcement learning methodologies that leverage human or AI feedback to improve LLM performance have recently been drawing attention. Existing RLHF (Reinforcement Learning from Human Feedback) goes through two separate stages—Reward Model training and Policy training—and this process has limitations in that training can become unstable or performance improvements can be difficult.

At this year's NeurIPS 2023, two key studies addressing these issues drew attention.

  1. DPO (Direct Preference Optimization): A methodology proposed by Stanford University that learns the policy directly from preference data without training a separate reward model. It transforms the complex process of existing RLHF into a single optimization problem, showing higher performance and more stable training behavior.

  2. Fine-Grained Human Feedback: A methodology proposed by the University of Washington that uses detailed feedback defined with clear metrics instead of simple preference information. This enables more effective training across various metrics, and it has demonstrated superior performance over existing methods on Detoxification and Long-Form QA tasks.

Going forward, research using such specific feedback signals to implement LLMs optimized for diverse user preferences is expected to accelerate further.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.