Contextual Answers Outperforms Foundation Models in Question-Answering Performance
Key point
AI21 Labs' Contextual Answers model surpassed leading Foundation Models, recording superior results in question-answering performance and reliability metrics.
Details
The three key metrics that determine the quality of a language model's question-answering (QA) is Output accuracy, Context integrity, and Answer relevance. In particular, when adopting enterprise GenAI solutions, minimizing Hallucination to ensure reliability is more important than anything else.
AI21 Labs conducted an experiment using the SQuAD 2.0 dataset to verify the performance of its QA-specialized model, Contextual Answers. The experiment was conducted on 1,000 samples with a 50:50 split of answerable versus unanswerable questions.
As a result of the experiment, Contextual Answers showed performance that surpassed the following major Foundation Models:
- Claude 3 Sonnet & Haiku
- GPT-4 Turbo & GPT 3.5 Turbo
- Mixtral 8x7B
In particular, Contextual Answers demonstrated overwhelming reliability in Context integrity, which accurately identifies questions with no answer to prevent hallucination, and Answer relevance, which delivers only the core content without unnecessary information. This shows that a Task-Specific Model optimized for a specific task can achieve higher efficiency in a particular domain than a general-purpose model.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.