LLM Hallucination Measurement Leaderboard Released
Key point
An open-source leaderboard for evaluating factuality and faithfulness hallucinations in LLMs has been released.
Details
The Hallucinations Leaderboard, designed to systematically measure Hallucination, a chronic problem in LLMs, has been released. This project aims to evaluate the reliability of various open-source models.
Hallucinations are evaluated by classifying them into two types:
- Factuality Hallucinations: Cases where the generated content contradicts actual facts.
- Faithfulness Hallucinations: Cases where the generated content does not align with the user's instructions or the given context.
The leaderboard evaluates models through zero-shot and few-shot in-context learning based on EleutherAI's LM Evaluation Harness. The related research can be found in a paper published on arXiv.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.