Open Text-to-Image Preference Dataset Released
Key point
The Hugging Face community has released an open-source preference dataset for training text-to-image generation models.
Details
Hugging Face's 'Data is Better Together' community has released the Apache 2.0-licensed text-to-image preference dataset (Open Image Preferences). The dataset aims to provide preference data that has been lacking for improving the performance of open-source image generation models.
The dataset covers a diverse range of model families and complex prompts, with the following features:
- Prompt diversity:
distilabelwas used to expand the categories and complexity of prompts with synthetic data, improving data quality. - Safety filtering: Multiple text- and image-based classifiers were used to remove NSFW and harmful content, with manual review to ensure safety.
- Ready-to-use resources: In addition to the preference dataset, a binarized version and the Flux-dev-lora-finetune model are also provided.
Users can download the dataset from the Hugging Face Hub, and the pre- and post-processing code is available on GitHub.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.