Key Insights
- Generative artificial intelligence has worsened a long-standing publish-or-perish problem in the academic community by making scientific paper fabrication trivially cheap and fast.
- AI detection systems catch what they can, but their creators openly admit that the race against improving AI models is one they may not win.
- Research integrity experts say the only lasting solution is changing how scientists are evaluated: away from volume and toward quality.
“Certainly, here is a possible introduction for your topic,” begins a retracted paper in the Elsevier journal Surfaces and Interfaces. This is one of ChatGPT’s standard openers when responding to a user prompt. The researcher submitted the manuscript beginning with this line, and it passed peer review. No one noticed the mistake until after publication.
Publishing is crucial for many scientists’ careers, especially those in academia, and a key part of securing jobs, funding, and tenure. The pressure to publish has always led some to take shortcuts. But with generative AI, a researcher can now create a convincing manuscript in just a few hours, even adding fake data and citations. Many journals don’t have the resources or motivation to catch this. As a result, generative AI has made an existing publish-or-perish problem worse.
“AI is an accelerant, not the kindling,” says Ivan Oransky, cofounder of Retraction Watch, which tracks retractions across scientific literature. “The kindling is publish or perish.”
A study published last July examining over 15 million PubMed biomedical abstracts worldwide found that at least 13.5% of papers published in 2024 showed signs of AI assistance. The rate varied by field, journal, and country—climbing as high as 40% in certain disciplines and running higher in non-English-speaking countries.
“AI is an accelerant, not the kindling. The kindling is publish or perish.”
“There are a lot of ways that AI can help out in scientific writing,” says David Resnik, a bioethicist at the National Institute of Environmental Health Sciences. It can smooth prose, improve grammar, and help non-native English speakers present work more clearly. “It can even draft sections of a paper with appropriate human oversight,” Resnik says. “But at the end of the day, the human being must be responsible, and the contribution of AI should be appropriately disclosed.”
Another study, published in August 2025, reported that the number of papers published by paper mills—organized operations that fabricate and sell fraudulent research—is doubling every 1.5 years, a rate much faster than the growth of legitimate scientific output.
The challenges posed by problematic AI-generated papers span a spectrum from the obviously absurd to the dangerously convincing. For example, at one end are manuscripts that in which signature phrases of generative AI such as “As an AI language model” left in the text because no one read the paper carefully enough to notice.
At the other end are papers that use AI to generate text, fabricate datasets, and create realistic-looking figures that evade standard detection. As Resnik points out, “Generative AI can make digital images completely from scratch. It’s really, really hard to tell the difference.” These papers look legitimate on the surface; they follow the conventions of scientific writing, cite references, and present data that conform to expected statistical distributions.
Once published, these papers spread. “It’s a problem for science as a whole,” Resnik says. Other scientists build on results that don’t exist, wasting time, funding, and resources. These flawed studies also skew funding decisions and steer research in the wrong direction. “Other researchers cite these papers, and AI systems train on them,” says Guillaume Cabanac, a research integrity expert and computer scientist at the University of Toulouse. Each time this flawed information is repeated or used, the error grows, ultimately threatening public knowledge.
All this undermines public trust in science, which is already fragile. It’s also unfair: those creating long publication records with the help of AI compete with honest researchers for jobs, grants, and recognition. Many researchers are asking whether the only way to combat the problem is systemic change—removing the reason researchers are being pushed to use AI in the first place.
Telling the truth about AI use
Most major publishers require authors to declare their use of AI. The journal Science has banned AI-generated text outright, the American Chemical Society (ACS publishes C&EN) requires a detailed acknowledgment, and the Committee on Publication Ethics has updated its guidance. To further address the issue, a coalition of organizations is developing a global AI disclosure standard.
“We don’t need generative AI to make a dog’s breakfast out of scientific literature.”
Resnik thinks disclosure needs teeth. He’s working on a paper proposing what he calls data attestation statements—a signed declaration submitted with every paper or grant application that the data are real and the original records exist, and that they will be shared if requested.
The idea is that real science can always be traced back to its source. “When you have AI, you don’t have an original source. It’s just made up,” Resnik explains. He also raises the possibility of digitally certifying certain experimental images the way fine art is certified—logging exactly which instrument produced an image and on which day, and who was logged in. If someone can prove a result came from a real microscope on a real Tuesday, fabrication becomes much harder.
However, the limitations of disclosure and data sharing remain. They depend on the author’s honesty in a system in which the incentive to conceal AI use is strong. This is where detection efforts become essential.
Detection systems may be losing the race
Cabanac built the Problematic Paper Screener in 2020 after spotting roughly 40 papers that showed definite signs of being produced by text generators but had somehow cleared peer review. He created a program to systematically screen scientific papers for signs of fabrication and misconduct and to publish the results online. Today, the screener trawls 160 million papers—and counting—indexed by the bibliographic database called Dimensions.
Phrases that are grammatically correct but scientifically nonsensical, like formic corrosive in place of formic acid and weighty metals instead of heavy metals, are what Cabanac calls tortured phrases. These are not scientific concepts, he says. That is exactly how he knows something is wrong.
For AI-generated text specifically, Cabanac looks for what he calls smoking guns: traces left by ChatGPT when users copy and paste its output without reading it. “Regenerate response” is one he sees often. “No researcher, no human would say that,” Cabanac says. “It is hideous.”
When the screener flags a paper, Cabanac posts a comment on PubPeer, a postpublication peer review platform, and contacts the publisher directly. More than 3,000 papers have been retracted as a result of his work.
Because Cabanac publishes his methods openly, similar tools have been widely adopted across scientific organizations and publications. Several major publishers have now implemented tortured-phrase screening at submission, catching problematic manuscripts before they reach peer review.
“I think [detection is] an arms race. . . it’s ultimately probably a losing battle.”
But detection remains a rearguard action, and its creators are honest about its limitations. Even Cabanac acknowledges that the tools only catch crude fraud. “I suspect there are cases where people generate datasets that are more difficult to find.”
Resnik’s 2025 paper says the potential for generative AI to create highly realistic fake data is “a ticking time bomb” in scientific publishing: there is no reliable method for distinguishing AI-fabricated datasets from real experimental data, and as models improve, there may never be one. He warns, “I think it’s an arms race. As tools improve, models get better at evading them, and users also get better at evading them. I think it’s ultimately probably a losing battle.”
“Even if we were to get a detector that works, it won’t solve all our problems,” Oransky says. “This notion that if we just tackle the AI problem we’ll get back to something great—it was never so great. Editors have been complaining about not being able to find peer reviewers for decades.”
Oransky sees AI use as almost inevitable, not just for publishing papers but even for grant applications. Applying for grants is so burdensome that “you should just accept that AI is going to be how people do most of it,” he says.
Targeting the root cause
Research integrity experts, along with most scientists, agree that the root cause of AI-driven and other fraud is the incentive structure of modern science. Researchers are evaluated primarily on the number of papers they publish and the prestige of the journals in which those papers appear. Tenure, funding decisions, promotions, salary, immigration status for international researchers, and institutional rankings all depend, to varying degrees, on publication volume. This system predates AI by decades and has been criticized for just as long. “We don’t need generative AI to make a dog’s breakfast out of scientific literature,” Oransky says. The only solution, he says, is to change what science rewards.
“If I had a magic wand,” Cabanac says, “we’d ask people to list only their top five articles, evaluating quality over quantity. In France, hiring and promotion committees are already asked to focus on a short selection of works.”
The San Francisco Declaration on Research Assessment, or DORA, was drafted in 2012 by a group of journal editors and publishers at the American Society for Cell Biology’s annual meeting. Now it is an independent worldwide initiative. It calls on universities, funders, and publishers to stop using publication volume and journal prestige as proxies for research quality. More than 25,000 institutions and researchers around the world have signed it. Canada’s federal funding agencies have introduced narrative curriculum vitae for grant applications. The US National Institutes of Health declared AI-generated grant applications ineligible last year and capped grant submissions at six per year per researcher.
“If I had a magic wand, we’d ask people to list only their top five articles, evaluating quality over quantity.”
Detection and disclosure buy time. “But only worrying about catching the bad guys isn’t going to get you anywhere,” Oransky says. “Should we end inequality, or should we punish people who end up having to steal?” We must do both things, he says. “You’re never going to let it all run rampant while you fix the upstream.”
“We’ve been talking about not measuring quantity over quality for decades,” Resnik says. “It hasn’t happened.” Meanwhile, it’s everybody’s problem, he adds. “Funders and funding agencies need to take responsibility, and Institutions need to develop better AI policies. And then of course, researchers should be responsible and disclose too. We need to slow down,” he says. “Publish less and let peer reviewers do their job.”
The barrier to fraud has never been lower. When fabricating a paper required months of effort and specialized skill, the cost was high enough to deter most people. When it requires an afternoon and a chatbot, the cost-benefit calculus changes entirely.