AI Briefing
KO

Mage

·2026.07.29 09:00

Key point

Microsoft has unveiled Mage, a family of efficient multimodal models built at a 4B parameter scale.

Details

Microsoft has unveiled Mage, a family of lightweight and research-friendly multimodal models built within a 4B parameter budget. Mage shares a 'codec-aligned efficiency' philosophy across both visual understanding and generation, designed to concentrate representational capacity where signal exists.

The model is compact enough to be trained, fine-tuned, and deployed even on modest hardware, while still delivering performance that competes with much larger open systems. The family is broadly divided into two core models.

  • Mage-VL: A codec-native streaming VLM for image and video understanding. It reads video in a codec-like manner (anchor/predicted frames, 16x16 patches) and uses a biologically inspired active streaming approach.
  • Mage-Flow: A native-resolution foundation model for image generation and editing. It is a 4B-scale generative stack combining Mage-VAE with a Native-Resolution MMDiT, supporting text-to-image generation and instruction-based editing, with Base, RL, and 4-step Turbo variants.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.