HuggingFace Releases IDEFICS
Key point
HuggingFace has released IDEFICS, an open-source reproduction of Flamingo, a multimodal model.
Details
HuggingFace has launched IDEFICS, an open-access visual language model (VLM) based on DeepMind's Flamingo. This model can take arbitrary sequences of interleaved text and images as input to generate text, and it has multimodal capabilities similar to GPT-4.
The key features are as follows:
- Two model scales: It offers Base and Instruct versions in 9B and 80B parameter scales.
- Data transparency: It was trained using public data such as Wikipedia and LAION, along with the 115B-token OBELICS dataset built directly by HuggingFace.
- Open-source orientation: It aims to increase transparency in AI research by reproducing a closed state-of-the-art model using only public data and models (LLaMA v1, OpenCLIP).
IDEFICS demonstrates excellent performance on tasks such as image question answering, describing visual content, and generating stories based on multiple images.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.