AI Briefing
KO

Sora System Card

·2024.12.09 09:00

Key point

OpenAI has released a system card detailing the technical architecture of its video generation model Sora and the measures taken to ensure its safety.

Details

Sora is OpenAI's video generation model that creates videos up to 1080p resolution from text, image, and video inputs. By combining a Diffusion model with a Transformer architecture, it delivers outstanding performance that maintains consistency even when a subject temporarily disappears from the frame.

Leveraging DALL·E 3's Recaptioning technology, it more precisely reflects users' text instructions, and it also features the ability to animate still images or extend existing videos.

Technically, Sora draws inspiration from the text token approach used in LLMs, compressing video into a low-dimensional latent space and then decomposing it into Spacetime patches for training. The dataset consists of the following:

  • Public data from industry-standard machine learning datasets and web crawling
  • Proprietary data through partnerships with Shutterstock, Pond5, and others
  • Human data including feedback from AI trainers and red teams

To prevent the risk of model misuse and generation of harmful content, strong filtering is applied starting from the pre-training stage. In addition, safety is continuously verified through red teaming activities conducted in collaboration with artists and experts from various countries.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.