AI Briefing
KO

Kakao Reveals Development Journey of Personalized Profile Image Generation Model 'Kollage Profile'

·2024.11.14 00:00

Key point

Implemented natural likeness using an HG-DPO-based backbone model and the Instantbooth module, validated via a photo booth at if(kakaoAI)2024.

1 / 9

Details

The Kakao Kanana Alpha organization revealed the development process of Kollage Profile, a model specialized for personalized profile image generation. This model adopts a diffusion approach, taking text and face images as conditional inputs to generate realistic images that are similar to the original person yet natural.

Core Technologies and Training Strategies

To improve model performance, a dedicated backbone model was fine-tuned using the HG-DPO (Human Generation Direct Preference Optimization) methodology. This extends DPO from the LLM field to image generation, leveraging large-scale datasets based on PickScore and Aesthetic score to improve the quality of person generation.

The Identity module was redesigned based on Instantbooth research. By separating and injecting high-level facial information (global embeddings) and pixel-level information (patch embeddings), it naturally reflects likeness in terms of shape, position, and fine details. During training, an Identity loss function that directly improves facial similarity was used in parallel with the diffusion loss function.

Datasets and Safety Measures

The training data consists of image-text pairs containing people, utilizing both Single Image of Single Identity (SISI) and Single Image of Multiple Identities (SIMI) data. Additionally, safety devices such as invisible watermarks, face forgery detection, and inappropriate word filtering were applied to prevent misuse like deepfakes.

Service Application and Future Plans

Currently, models in two sizes, Nano and Essence, have been completed and are planned for deployment optimized for each service based on generation speed and resolution. At the if(kakaoAI)2024 event, a demo successfully generated profile images in four styles, including animation and crayon, in about 30 seconds via the 'Kanana Photo Booth'. In the future, personalization technology is planned to be applied to the video generation model Kinema by Kanana as well.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.