Moderation API Upgraded with New Multimodal Moderation Model
Key point
OpenAI has introduced a new GPT-4o-based multimodal moderation model into the Moderation API.
Details
OpenAI has introduced omni-moderation-latest, a new moderation model based on GPT-4o that can detect both text and images, into the Moderation API. This model is more accurate than the previous model, and shows especially strong performance in non-English languages.
Key improvements are as follows:
- Multimodal harm classification: It can combine images and text to assess harm across 6 categories, including violence, self-harm, and sexual content. Multimodal support is currently limited to these categories, with plans to expand coverage in the future.
- New text-only categories: The
illicitcategory, which covers instructions for illicit activity, and theillicit/violentcategory, which involves violence, have been added. - Improved accuracy: In tests across 40 languages, internal multilingual evaluations showed a 42% performance improvement, with dramatic gains in low-resource languages such as Telugu (6.4x) and Bengali (5.6x).
In addition, probability scores have been adjusted to allow finer control over moderation decisions. This new model is available free to all developers through the Moderation API.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.