On the Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
Key point
RL-finetuned VLMs show improved visual reasoning performance, but exhibit vulnerability to textual perturbations.
Details
Reinforcement Learning(RL) fine-tuning has established itself as a key technique for boosting LLM performance on reasoning-centric tasks, and its scope has recently been expanding to VLM(Vision Language Model)s as well.
RL-finetuned VLMs show performance improvements on visual reasoning benchmarks, but they still suffer from weak Visual Grounding, Hallucination, and excessive reliance on textual cues.
According to the research findings, even simple textual perturbations—such as incorrect captions or inaccurate Chain-of-Thought(CoT) reasoning paths—significantly degrade the model's robustness and reliability. In particular, this phenomenon becomes more pronounced as CoT consistency decreases.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.