Lance: Unified Multimodal Modeling through Multitask Synergy
·2026.05.25 09:00
Key point
ByteDance has unveiled Lance, a 3B-scale unified multimodal model that supports understanding, generation, and editing of images and videos.
Details
Lance is a lightweight native unified multimodal model that supports understanding, generation, and editing of images and videos within a single framework.
With only 3B active parameters, it delivers strong performance on image generation, editing, and video generation benchmarks.
Key features are as follows:
- Multitask Synergy: Performs diverse multimodal tasks in an integrated manner within a single model.
- Efficient Training: Trained from scratch within a budget of 128 A100 GPUs through a staged multitask recipe.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.