AI Briefing
KO

STARFlow2 Achieves Unified Multimodal Generation by Connecting Language Models and Normalizing Flows

·2026.08.25 09:00

Key point

Apple released STARFlow2, which combines LLMs and normalizing flows to generate text and images in a unified manner.

Details

Existing unified multimodal models suffer from structural drawbacks, such as sacrificing visual fidelity, harboring structural asymmetry, and degrading pretrained understanding capabilities.

Apple researchers noted that Autoregressive Normalizing Flows share the same causal mask, KV-cache mechanism, and left-to-right structure as LLMs. Based on this, they proposed STARFlow2.

STARFlow2 is based on the Pretzel architecture, vertically interleaving a frozen pretrained VLM stream and a TARFlow stream via residual skip connections. Both streams operate under the same causal mask, preserving pretrained multimodal understanding while enabling high-fidelity continuous image generation.

The Deep-Shallow Flow design combined with an integrated FAE latent space supports cache-friendly interleaved generation, where both text and visual outputs enter the KV-cache directly without re-encoding. Experimental results demonstrated strong performance on image generation and multimodal understanding benchmarks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.