Summarize API Outperforms OpenAI Models
Key point
AI21's Summarize API produced more accurate and condensed summaries than OpenAI models.
Details
AI21 Labs compared its Summarize API against OpenAI's Davinci-003 and GPT-3.5-Turbo across various prompts, academic data, and real-world data. The automatic metrics were faithfulness rate and compression rate, and human experts determined pass rate through blind evaluation.
The results showed that the Summarize API was overall better than or at least comparable to the OpenAI models. In particular, on real-world data it showed fewer hallucination and reasoning violation cases, and OpenAI summaries were classified as 'Very Bad' at a rate 2 to 4 times higher. GPT-4 was not yet included at the time of the comparison.
Key figures are as follows.
- 18% improvement in pass rate over Davinci-003
- 19% improvement in faithfulness score over Davinci-003
- 27% improvement in compression rate over GPT-3.5-Turbo
- Standard deviation of compression rate was at least 50% lower, resulting in more consistent summary lengths
The Summarize API is a summarization-dedicated model that works immediately with just the source text input, without any separate prompting, making it simple to use and more consistent in its results. In contrast, OpenAI models showed no meaningful improvement even with strong prompt engineering applied.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.