AI Briefing
KOSign in

Google and DeepMind researchers propose CO₂Jump for consistent text-image generation

·2026.09.30 16:28

Key point

The new sampler requires no additional training and uses one model forward pass per denoising step to revise low-confidence tokens.

Details

Researchers from Google, Google DeepMind, and Stony Brook University introduced CO₂Jump, a sampler designed to resolve mismatches in joint text and image generation. The method addresses the issue where a model might describe a correct solution (e.g., a maze path) while generating a visually inconsistent image.

How It Works

CO₂Jump uses text confidence and cross-modal attention to guide image updates during the sampling process. It allows low-confidence tokens to be masked and regenerated, enabling the model to revise earlier decisions as generation progresses. The sampler operates with one model forward pass per denoising step and requires no additional training, relying instead on task-specific fine-tuned models.

Evaluation and Results

The team evaluated CO₂Jump on image editing, maze solving, and nonograms, introducing three new datasets: JEdit-1M, JMaze-200K, and JNono-200K. On puzzle benchmarks, success required both the textual answer and the generated image to be correct. Across 8–512 sampling steps, CO₂Jump was the only sampler compared that showed monotonic improvement in both editing quality and grounding.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.