AI Briefing
KO

Robostral Navigate: AI Navigation Based on a Single Camera (5-min read)

·2026.07.09 09:00

Key point

Robostral Navigate is an 8B-scale embodied AI model that autonomously navigates complex environments using only a single RGB camera.

Details

Robostral Navigate is an 8B model that autonomously navigates complex environments using only a single RGB camera, without LiDAR or Depth sensors. This model achieved a 76.6% success rate on the unseen dataset of the R2R-CE (Room-to-Room in Continuous Environments) benchmark, demonstrating higher efficiency and performance than existing multi-sensor-based approaches.

This model adopts a Pointing-based navigation approach that predicts target coordinates within the camera's field of view and the heading direction upon arrival. This allows it to flexibly respond to camera characteristics or changes in world scale, and when the target is outside the field of view, it compensates by switching to local coordinate system-based movement commands.

Key features are as follows:

  • High versatility: Applicable to various robot forms including wheeled, legged, and flying robots, and works regardless of robot size.
  • Efficient training: Trained on approximately 400,000 trajectories collected from 6,000 scenes in a simulation environment.
  • Prefix-caching technology: Uses a tree-based Attention-masking strategy to compress entire episodes into a single sequence, maximizing training efficiency.
  • Self-developed model: Built on a proprietary Vision-Language Model specialized for Grounding tasks, rather than relying on existing open-source VLMs.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.