AI Briefing
Sign in

Deriving Policy Gradients for LLMs: From REINFORCE to GRPO and Handling Inference Mismatch

·2026.09.27 16:00

Key point

This article derives policy gradient methods for LLMs, explaining how REINFORCE evolves into GRPO through baseline variance reduction and addressing the gradient bias caused by inference engine mismatches.

1 / 5

Details

From REINFORCE to Policy Gradients

The text derives the policy gradient estimator for LLMs, starting with the REINFORCE algorithm. It explains that the gradient of the objective function can be estimated by sampling completions from the policy and weighting the score function ($\nabla_\theta \log p_\theta(y)$) by the reward. This is described as "supervised fine-tuning on your own samples, weighted by reward." The article highlights that while REINFORCE is unbiased, it suffers from high variance.

Baselines and GRPO

To reduce variance, the article introduces the concept of subtracting a baseline $b$ from the reward, creating an advantage $A = R - b$. It notes that this does not change the expected gradient but significantly reduces variance. The article specifically discusses GRPO (Group Relative Policy Optimization), which uses the mean reward of a group of $G$ completions as a critic-free baseline. This approach replaces the need for a separate value network (critic) used in PPO, saving computational resources. The text also mentions that GRPO typically normalizes by the group's standard deviation, though the derivation focuses on mean-centering.

Implementation Challenges: Inference Mismatch

A critical practical issue addressed is the discrepancy between the training model and the inference engine (e.g., vLLM or SGLang) used to generate completions. The theoretical derivations assume samples are drawn from the same distribution $p_\theta$ being differentiated. However, in practice, numerical precision differences or asynchronous weight updates cause the inference engine to use a slightly different distribution. This breaks the zero-mean property of the score function, potentially introducing bias into the gradient estimates when baselines are applied.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.