A new benchmark spots AI-written papers by their broken reasoning

Researchers at Seoul National University and the University of Minnesota posted a preprint to arXiv on 30 September proposing a different way to detect AI-generated academic papers. Rather than looking at how the prose reads, which is what token-level detectors do, they test whether the scientific reasoning holds together across the paper. Their framing is that each section looks plausible on its own while the connections between them break down.

The benchmark, SciSlopBench, pairs 390 AI-generated papers with human-written ones matched by research problem and contribution type, mostly in computer science but reaching into the life, social and natural sciences. Six measures grouped under structure, argument and artifacts identify the AI paper in each pair with 85.9% accuracy, against 68.7% for Binoculars, an existing detector.

The result that travels furthest is not the detection score. Higher scientific slop scores accompany lower ICLR review ratings, and separate rejected from accepted papers above chance in every year from 2017 to 2025, which suggests the measures are picking up something human reviewers already penalise.

Treat the numbers as the authors’ own and not yet peer-reviewed. The paper also reports that optimising directly against its measures triggers reward hacking, and that its mitigation framework closes 63% of the remaining gap over the strongest revision baseline.

Source: the preprint on arXiv.


Related