AI Briefing
KO

LLM Evaluation Volatility Based on Prompt Format

·2024.04.30 09:00

Key point

A study has found that minor formatting changes in prompts significantly affect LLM benchmark scores and model rankings.

1 / 2

Details

Hugging Face's Leaderboards and Evals research team released experimental results showing that LLM benchmark performance is highly sensitive to minor formatting changes in prompts.

Even when the same amount of information is provided, model performance varies significantly depending on question tags, the way choices are presented (A, B, C, D or 1, 2, 3, 4), and the method of log-probability calculation. In experiments using the MMLU benchmark, performance differences of around 10 points or more occurred depending on the model, and even a phenomenon where rankings between models were reversed was observed.

This phenomenon poses a major obstacle to objectively comparing the actual performance of models. To address this, Hugging Face is collaborating with Dottxt to research Structured Generations approaches that induce consistent outputs regardless of prompt format.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.