AI Briefing
KOSign in

SemanTok: Predictable Semantic Tokens Enable Efficient Autoregressive Video Generation

·2026.10.07 23:56

Key point

A 201M-parameter SemanTok AR model matches or outperforms a VideoFlexTok AR model 3.4x its size.

Details

Video-based world models increasingly combine the scalability of autoregressive (AR) prediction with the visual quality of diffusion models, making the choice of scene tokenizer critical for both fidelity and semantics. SemanTok is a new flexible video tokenizer designed to address limitations in existing methods, which often apply representation-alignment (REPA) losses only to early decoder hidden states—a target the decoder can partially satisfy from its noised input alone.

How SemanTok Works

SemanTok feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct these features from each retained token prefix. This approach ensures that the first coarse tokens carry the clip's global semantics, while later tokens specify finer details. By aligning semantics at every token prefix, the model achieves high semantic alignment and video fidelity across all AR model sizes.

Performance and Efficiency

The method demonstrates significant efficiency gains:

  • A 201M SemanTok AR model matches or beats a VideoFlexTok AR model that is 3.4x its size.
  • Larger SemanTok AR models continue to improve fidelity.
  • Short token prefixes are cheaper to predict and yield better generation fidelity, as pixel detail is deferred to later tokens.

Robustness

SemanTok maintains semantic alignment on out-of-distribution classes and provides the decoder with higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation tasks.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.