AI Briefing
KO

Kakao Brain releases open-source ViT and ALIGN

·2023.03.06 09:00

Key point

Kakao Brain has open-sourced the COYO dataset, containing 700 million image-text pairs, along with ViT and ALIGN models trained on it.

1 / 2

Details

Kakao Brain has released the open-source dataset COYO, composed of 700 million image-text pairs, along with ViT and ALIGN vision-language models trained on it. This marks the first case of an ALIGN model being released for free and open-source.

These models use the same architecture as Google's existing ViT and ALIGN models, but were trained on the publicly released COYO dataset, allowing researchers to access the dataset and reproduce the models.

Key features and performance:

  • Performance: The ALIGN-B7-Base model achieved comparable performance to Google's model on Image KNN classification tasks despite using less data (700 million vs 1.8 billion), and recorded superior performance on the MS-COCO retrieval task.
  • COYO Dataset: Contains 700 million image-text pairs and provides rich metadata such as Aesthetic Score, watermark score, and face count, enabling finer-grained control compared to LAION 2B.
  • Accessibility: ViT and ALIGN demos can be experienced directly through Hugging Face.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.