NeurIPS Rejects 178 Papers Using AI Detector
Key point
NeurIPS rejected 178 papers using the Pangram AI detector, but reliability concerns emerged as the chairs' own papers were judged to be up to 69% AI-generated.
Details
The NeurIPS Position Paper Track used the exclusive AI detection tool Pangram to reject 178 papers, representing 18.4% of all submissions, without human review or an appeal process. This process revealed serious flaws in the detector's reliability and fairness.
Detector Malfunctions and Reliability Issues
Independent researchers ran the recent papers of the three track chairs through the same detector, resulting in judgments of 24% to 69% AI generation. This implies that the chairs themselves risked rejection under their own rules. Pangram's default settings classified 42.7% of all submitted papers as 90-100% AI, and the text window was reduced to lower this figure, adjusting the final rejection rate to 12.7%.
Disadvantage for Non-Native English Speakers and Procedural Flaws
According to Stanford research, 61.22% of human-written essays by non-native English speakers (ESL) are misidentified as AI. NeurIPS did not disclose such demographic calibration data, creating a structural bias disadvantageous to non-native English-speaking researchers. Additionally, 22 papers with detector scores above 0.5 where authors denied AI use were used to prove the authors' dishonesty based on black-box scores.
Follow-up Actions
Rejected papers do not remain as records of misconduct and can be resubmitted to ICLR (deadline September 25) or ICML, among others.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.