AI Briefing
KO

A Verification Hole, Oral

·2026.04.15 15:12

Key point

The limitations of natural language metrics in SQL code generation evaluation and a **20% false positive** rate have been pointed out.

Details

A paper selected as an ICLR 2025 Oral has been criticized for using a natural language metric instead of an execution metric to evaluate SQL code generation.

The author says that after re-examining this evaluation method, they confirmed an approximately 20% false positive rate in their own tests.

The key issues are as follows.

  • Choice of evaluation metric: relies on a natural language-based metric rather than actual execution results
  • Verification reliability: tests showed a false positive rate of around 20%
  • Appropriateness of selection: raises the question of why the paper received an Oral despite this flaw

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.