Artificial Intelligence & Computing — AI Detection

How One Team Scored 12,750 Academic Papers and Found a Third Written by AI

A careful measurement of AI-generated writing across arXiv preprints finds about a third of recent papers read as machine-written. The headline number is real, but the real story is the 20% of genuinely human papers the detector gets wrong — and the honest limits of any score.

Headlines of the form “N percent of X is now AI” are common, and most are not worth reading, because the tool behind the number also marks some genuine human writing as artificial. The team at UnSlop designed their study around exactly that objection. Their detector is calibrated specifically for academic writing: at a 0.4% false-positive rate it clears 99.6% of real pre-LLM scientific text and recovers 85% of AI-written academic text. They set the threshold using papers submitted in 2021 and 2022, before ChatGPT, treating them as ground-truth human and choosing a line that trips exactly 0.4% of them. Every number they report sits above that floor.

They scored the full body text of 12,750 papers across ten field groups, pulling only the first version of each PDF so that a paper revised in 2026 could not leak modern phrasing back into its 2023 slot. Abstracts were ignored because they understate the signal — the same paper can score under 20% on its abstract and over 70% on its body. The result is a clean curve: flat at 0.4% through 2021 and 2022, lifting within months of ChatGPT, and climbing in two waves to about 32% over the most recent complete quarter, peaking near 39% in early 2026.

Interpreting the spread across fields requires honesty. Mathematics’ near-zero figure is not evidence that mathematicians are writing everything by hand; math papers are dominated by notation and theorem-proof structure, and once equations are stripped away, the remaining prose is too sparse and too unlike the scientific English the detector was trained on. A low score can mean a detector blind spot as easily as it can mean genuine human authorship. A flag also does not prove authorship — it indicates text that reads as machine-written, which includes heavily AI-edited drafts. The measurement is useful as a lower bound and a warning, not as a verdict.