Tencent Unveils RxBrain, a Multimodal Robotic Cognition Model
·2026.07.15 18:30
Key point
Tencent has unveiled the RxBrain model, which combines language reasoning and visual imagination to plan robot actions.
Details
Tencent has released RxBrain-1.0, a unified multimodal foundation model for Embodied Cognition, on Hugging Face. This model is designed to combine language reasoning and visual imagination so that robots can understand the physical world and plan actions.
Key Features:
- Embodied Understanding & Reasoning: Performs question answering and Chain-of-Thought on images and multi-frame videos
- World State Prediction: Visually predicts future frames that will appear in the physical world when a specific action is taken
- Joint Subgoal Planning: Decomposes tasks into step-by-step subgoals, simultaneously generating the language-based action instructions and goal images to be reached required for each step
Technical Characteristics:
- Unified Mixture-of-Transformers (MoT): Uses a backbone of approximately 6.2B parameters, processing text, vision, and generation modalities together within a single autoregressive model.
- Flow-Matching Image Head: Implements text-to-image generation and multi-frame world model rollout through a flow-matching head that decodes the FLUX VAE latent space.
- Interleaved Reasoning + Imagination: Generates text reasoning and generated frames interleaved in a single sequence, combining symbolic planning with visual goals.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.