GPT-4o System Card
Key point
OpenAI released a system card detailing GPT-4o's multimodal capabilities and audio-related safety evaluations.
Details
GPT-4o is an autoregressive omni model that accepts text, audio, image, and video as input and outputs text, audio, and image. It processes all inputs and outputs through the same neural network, and supports real-time interaction at a human-like conversational level with an average response speed of 320ms.
This system card particularly focuses on the new risk factors that the newly introduced audio capabilities may pose. The key evaluation and mitigation targets are as follows:
- Speaker identification and unauthorized voice generation
- Potential generation of copyrighted content
- Ungrounded inference and generation of disallowed content
According to the evaluation results, GPT-4o's voice modality was not found to significantly increase Preparedness risk. Of the 4 Preparedness Framework categories, 3 were rated 'low,' while persuasion remained at the 'medium' borderline.
The model was trained on data up to October 2023. The dataset consists of the following:
- Web data: Public web page data for learning diverse perspectives and topics
- Code and math: Structured data to improve logical reasoning ability
- Multimodal data: Data for learning non-text input/output through images, audio, and video
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.