AI Briefing
KO

Show HN: Lance - image/video generation and understanding in a single model

·2026.05.21 06:01

Key point

ByteDance has released Lance, a 3B model that unifies image/video generation, understanding, and editing.

1 / 2

Details

ByteDance has released Lance.

It is a unified multimodal model with 3B active parameters that handles understanding, generation, editing of images and video within a single framework.

  • Trained with a 128-A100 GPU budget.
  • The transformer backbone was trained from scratch, excluding the ViT and VAE encoder.
  • Using a staged multi-task recipe, they released demos of text-to-video, video editing, multi-turn consistency editing, intelligent video generation, text-to-image, image editing, and image understanding.
  • They also presented strengths on image generation, image editing, video generation, and video understanding benchmarks.

A homepage, arXiv paper, and Hugging Face model page were released together.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.