AI Briefing
KO

AI Agents' Citation Errors: 15.9% Cite Inappropriate Sources

·2026.06.26 19:19

Key point

A new benchmark, OpenBioRQ, points out that 15.9% of papers cited by AI agents do not match the content of their claims.

Details

A new benchmark called OpenBioRQ evaluated the citation accuracy of AI agents using 12,553 unresolved biomedical research questions across 12 fields.

The study found that hallucinations where agents fabricate URLs themselves were rare (over 99% of URLs were valid), but the rate at which cited papers did not match the actual claims reached about 15.9%. This suggests that existing evaluation methods, which only check URL validity, may overestimate the reliability of agents.

Key observations are as follows:

  • Performance gap by model: On the high-difficulty question set, Gemini-3-Pro, Opus-4.7, and GPT-5.5 showed resolution rates of 29-60%, while open-weight models reached only about 17%.
  • Tool-use dropout phenomenon: As question difficulty increased, a behavioral collapse was observed in which agents stopped using retrieval tools altogether.
  • Importance of evaluation method: It was confirmed that existing URL validity checking methods make it difficult to measure an agent's actual logical citation accuracy.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.