RelateAnything Supports 17.2ms Inference on RTX 3080 Laptop
Key point
RelateAnything supports an inference speed of 17.2ms per image on the RTX 3080 Laptop GPU using TensorRT FP32 and has released the RA-4M training corpus, which includes over 474,000 images and 4.3 million relationships.
Details
Deployment and dataset information for the newly released open-source relationship prediction model RelateAnything has been made public. The model supports various deployment paths, including PyTorch (CPU/CUDA), TensorRT 10, ONNX, OpenVINO, and a browser demo (WebGPU). According to the README, the ViT-S/16+ model takes 17.2ms per image on the RTX 3080 Laptop GPU using TensorRT FP32 (batch size 1, 35 relationship types, including preprocessing and decoding).
The RA-4M corpus, used for training, consists of over 474,000 images, 4.3 million relationships, and 10,102 free-text predicates. This dataset was generated by having vision-language models write relationships based on numbered box markers, followed by filtering through deterministic geometric checks. The licenses are as follows: code under Apache-2.0, model weights under the DINOv3 license, and RA-4M annotations and masks under CC BY-NC 4.0.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.