Along Belfast’s Ormeau embankment, a mural asks “How can quantum gravity help explain the origin of the universe?” Like much of the city’s gable-end artwork, this tends to fade into the background somewhat. Even as a historian of science, what I know about quantum gravity could plausibly fit on the back of an envelope, not least because it largely comes from a hoax conducted 30 years ago.
In 1996 a physicist at New York University sent a paper called Transgressing the Boundaries: Towards a Transformative Hermeneutics of Quantum Gravity to the cultural studies journal Social Text, which published it in a special issue on the science wars. The paper was written to flatter the editors’ intellectual and political assumptions while making extravagant and untenable claims about physics. Its author, Alan Sokal, then announced that he had written nonsense on purpose to demonstrate the lack of rigour in social science publishing.
Social Text had no peer review, so the hoax demonstrated the credulity of two editors rather than a failure of scholarly gatekeeping. But Sokal raised an important point. Publication in an academic journal is generally taken to confer a certain amount of credibility to an argument or the conclusions reached from research. Career progression for researchers largely depends on accumulating many such publications. The algorithms used by search engines, social media and, increasingly, large language models, generally weight academic publications highly.
There is therefore an uneasy relationship between quantity and quality on both the demand and supply sides of academic publishing. This is illustrated most clearly by paper mills: businesses that manufacture studies, fabricate or recycle data, and sell completed papers or authorship positions to researchers who need publications.
How to change career: Advice from three people who have done it
Broadcaster Sarah Mulkerrins: ‘When I left Ireland, I didn’t think I’d ever get to work on GAA, so it’s a bonus’
Martin and Roman Kemp explore what Ireland does best: Sarcasm, breakfast, wondering what’s in the will
Women continue to bear the brunt of bad tech decisions
A paper in the BMJ this year published a machine learning screen of cancer literature. The researchers trained a BERT language model on 2,202 cancer papers that had been retracted and identified as paper mill products, matched with 2,202 papers treated as legitimate controls. BERT’s job was to learn whether the titles and abstracts of suspect papers shared recurring patterns.
[ The credibility crisis in scienceOpens in new window ]
The team then applied it to 2.6 million original cancer research articles from 1999-2024. It flagged 261,245 papers, just under 10 per cent, as similar to known paper-mill publications. The proportion rose from around 1 per cent in the early 2000s to more than 16 per cent in 2022. Flagged papers appeared across major publishers, not just marginal journals.
There are, of course, some caveats. The controls weren’t random; they used papers from Taiwanese and Scandinavian teams and from selected high-impact journals as proxies for reliable work. Further, the BERT tool only scanned titles and abstracts. Most importantly, having characteristics associated with a paper mill is not necessarily the same as being from a paper mill. Nevertheless, as the great and ever-relevant webcomic xkcd reminds us, correlation doesn’t imply causation, but it does waggle its eyebrows suggestively and gesture furtively while mouthing “look over there”.
In April, Reese Richardson, Spencer Hong, and Anna Abalkina released a data set, BuyTheBy, that compiles 18,710 advertisements from paper mills in India, Iraq and the former Soviet bloc. A first author credit costs from $57 to $5,631 (€49-€4,877), with a median around $788 (€683). This is clearly a segmented market with price discrimination; it supplies academic reputation to a clientele whose ultimate customer is a promotion committee counting their outputs.
Generative AI has changed the economics of this activity. It can produce fluent abstracts, vary recycled prose, generate plausible literature reviews and reduce the obvious linguistic irregularities on which earlier detection methods relied. Fabricated material need no longer be copied directly from an existing paper, making conventional plagiarism checks less effective.
The danger isn’t just that AI can write nonsense. It can write plausible, disciplined, and technically ornamented nonsense at a scale that human reviewers cannot readily digest.
It took me around 20 minutes to read the BMJ paper, and another 10 to wrap my head around the BERT model and methodology. At that rate, it would take me 541,667 days to read 2.6 million articles. Even if I could reliably spot the fakes, cite the right authorities and reach conclusions that reviewers expect to see, if I started a countdown today, it would hit zero in 3509. The only plausible way to do it is with a machine tool.
And therein lies the issue. Where careers, promotion, or professional accreditation depend on publishing, a market for manufactured publication is a predictable outcome. Generative AI may lower costs and enlarge that market, but it didn’t create the incentives. Unless academic reputation is tied more closely to reproducibility, data transparency or negative findings, we will use AI to screen for AI in an increasing number of papers that people are decreasingly likely to read.
Stuart Mathieson is research manager with InterTradeIreland















