AI Briefing
KO

Conditions for LLMs to Replace Humans in A/B Testing

·2026.08.14 03:57

Key point

Using LLM predictions as a proxy metric for A/B testing requires specific statistical assumptions and calibration procedures.

1 / 2

Details

Using LLMs to run experiments instead of actual users can drastically reduce time and costs. However, for LLM predictions to fully replace actual user responses, it must be statistically guaranteed that the experiment can identify the treatment effect of interest.

Experiments using the Upworthy dataset showed that the raw predictions of gpt-4o-mini recovered only 39% of the actual human treatment effect. This indicates that LLM results are not merely noisy but possess a systematic bias that converges the treatment effect toward zero. Such bias can lead organizations to underestimate product value and make incorrect decisions.

For LLM predictions to serve as valid surrogate metrics, the following two conditions must be met:

  • Surrogacy: The LLM's output must fully mediate the process by which the treatment effect leads to human outcomes. In other words, the LLM must capture all key factors influencing human responses.
  • Comparability: The calibration function describing the relationship between LLM predictions and human outcomes must remain consistent with past data in new experiments.

In conclusion, LLM-based A/B testing is possible, but it relies on specific assumptions rather than design guarantees, and its reliability decreases as new treatments diverge from past experiments.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.