LG AI Research 211
Key point
From the emergence of DALL-E through VQ-VAE-based models to LG AI Research's L-Verse, we look at the evolution of Vision-Language models.
Details
Since the emergence of GPT-3 in 2020, Large Scale Model Training has become the core of the NLP field. In January 2020, OpenAI unveiled DALL-E, a model with 120 billion parameters, opening the era of Zero-shot text-to-image generation that creates images from text.
DALL-E uses CLIP to analyze the similarity between images and text, and converts images into tokens in the same form as text via VQ-VAE, which are then trained within a Transformer structure. This involves a process of efficiently compressing and restoring image data.
Subsequent research has developed in the direction of achieving high performance with less data and fewer parameters. In particular, L-Verse, announced by LG AI Research, demonstrated outstanding performance capable of both text-to-image and image-to-text while using far fewer resources than DALL-E.
Recently, to address the information loss problem of the existing VQ-VAE approach, the paradigm is shifting toward new research on Large Vision-Language Models (LVLM) that combine DDPM and CLIP.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.