11 hours ago · 15 min read3080 words · Tech · hide · 0 comments

Earlier this year, the program chairs for one of the largest machine-learning conferences found references in submitted papers to publications that did not exist. This was not a thought experiment about what a language model might do. It was not a benchmark in which researchers asked a model to write a literature review and then counted the invented citations. These were papers submitted to ICLR 2026, accompanied by bibliographies that were supposed to describe prior work. The conference built a screening system, sent the flagged references to area chairs, and then had the program chairs check them again. Every paper with a confirmed hallucinated reference was desk rejected. The most interesting part of the program chairs’ account is how much work it took to establish that a paper in a bibliography was not a paper. The automated system produced false positives. Translated titles looked suspicious. At least three people reviewed every confirmed case. The conference did not trust one…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.