AI Briefing
KO

Stability AI Annual Integrity Transparency Report

·2025.09.18 02:26

Key point

Stability AI has released its annual integrity transparency report to demonstrate the safety and transparency of its model development and deployment process.

Details

Stability AI released its Annual Integrity Transparency Report, built on Safety-by-Design principles, to responsibly build and deploy generative AI. The report aims to build trust with users, developers, and policymakers by transparently disclosing the design, testing, monitoring, and misuse-prevention framework of its models.

Training data draws on publicly available internet data, third-party partner data, and Synthetic Data generated by researchers. Sources where harmful content circulates, such as the dark web or adult websites, are excluded, and data is filtered through in-house-developed and open-source NSFW classifiers. To date, CSAM (Child Sexual Abuse Material) has been confirmed at 0% within the training datasets.

The following multi-layered mitigation measures are applied to ensure the safety of models and platforms:

  • API level: Real-time content filters and classifiers detect policy-violating inputs and outputs, and a CSAM hashing system is integrated to block and report harmful material.
  • Model level: Models are deployed using Fine-tuning and Safety LoRA techniques based on insights gained from Red Teaming.

Red Teaming is a key process for identifying serious risks in models. It involves both internal and external experts, and includes ongoing risk assessments such as collaborating with OCCIT, a UK law enforcement agency, to verify the safety of the Stable Diffusion 3 model.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.