AI Briefing
KO

Mistral Moderation API

·2024.11.07 09:00

Key point

Mistral AI has released an LLM-based Moderation API that can classify content across 9 categories.

1 / 2

Details

Mistral AI has released the same content moderation API used in its own service, Le Chat. Users can customize this tool to fit their own applications and safety standards.

This model is an LLM classifier that classifies text input into 9 categories, and it provides two endpoints.

  • Raw text: for classifying plain text
  • Conversational content: for classifying conversational content by evaluating the last message within a conversation's context

The model is a natively multilingual model supporting 11 languages, including Korean. In particular, it focuses on preventing model-generated harms such as inappropriate advice or exposure of personal information (PII).

Performance was verified using the AUC PR metric on an internal test set, and Mistral AI plans to continue collaborating with the research community to advance safety technology.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.