AI can now produce a beautifully structured peer review in seconds. Give it a manuscript and it will identify methodological problems, question the statistics, flag inconsistencies, suggest analyses, recommend missing references and produce an intimidating list of major and minor comments in immaculate academic English. In many cases, the resulting review may look more comprehensive than the one you would receive from a human expert who spent a significant amount of time reading the paper.
So perhaps we should ask the obvious question: do we still need human reviewers?
I am not suggesting that journals sack their reviewer databases and hand everything over to ChatGPT. The more useful question is what, exactly, we still need the human reviewer for. If AI can perform much of the visible labour of peer review faster than humans, asking a senior academic to spend an evening checking references and table formatting increasingly looks inefficient.
The machine can do much of the checking. What it cannot reliably do is decide what matters.
Reviewers are already using AI
This is not a hypothetical future. A 2025 Frontiers survey of 1,645 active researchers reported that 53% of peer reviewers were already using AI tools during peer review.
The question is therefore no longer whether reviewers use AI. It is whether they use it as an assistant or allow it to substitute for judgement.
My Personal Experience
I recently had reason to wonder about that distinction. An article I had submitted to a reputed journal remained under review for close to a year despite repeated reminders. Eventually, a decision arrived based on a single reviewer report.
The report was about two pages long, neatly organised with headings, subheadings and bullet points.
To me, there was little doubt that it had been generated, or at least largely drafted, by an LLM. The style was simply too familiar. I could not prove it, so I asked the editor whether AI had been used. I was reassured that the reviewer was very much human—and apparently a very senior one.
That reassurance worried me more.
The problem was not whether the reviewer existed. It was whether the report reflected the reviewer’s own judgement. There were plenty of comments, but little engagement with the central issues. It looked comprehensive, but it did not feel thoughtful.
That is where AI-assisted peer review can go wrong: an impressive-looking review can create the appearance of expert scrutiny without containing much expert judgement.
Peer review is not primarily an exercise in finding faults or saying the most!
It is an exercise in deciding which faults matter; what actually changes the interpretation, importance or publishability of the work.
When AI does not want you to reject…
A reviewer who is very good at finding things to improve but rarely concludes that a paper should be rejected is of limited value. If the default response (this is what AI does) is always “revise and improve”, peer review loses one of its central functions: deciding when a paper is not sufficiently novel, important or methodologically sound to justify publication.
There is evidence that LLM reviewers may have exactly this tendency. Janis Keuper generated 1,000 reviews of 2024 ICLR papers using several LLMs and found that many models recommended acceptance in more than 95% of cases, even without any attempt to manipulate them. These were machine-learning papers, so the finding cannot simply be extrapolated to biomedical publishing, but the degree of leniency is striking.
Russo and colleagues found a related pattern in actual ICLR 2024 reviews. They estimated that at least 15.8% had received AI assistance. Among papers close to the acceptance threshold, receiving an AI-assisted review was associated with a 4.9-percentage-point higher acceptance rate, suggesting that AI assistance may influence the decision itself, rather than merely improve the wording of a review.
Why might this happen? An LLM is very good at finding problems and suggesting how they could be fixed. But deciding that a paper should not be published is a different task. It requires judging whether the central idea is sufficiently novel or important, whether a flaw undermines the whole study, and whether further revision is actually worthwhile.
The same problem may affect authors.
In my own use, I have repeatedly noticed how readily an LLM describes a manuscript as promising before offering an extensive list of improvements. I could find no direct evidence that this systematically gives authors false confidence, so this remains an observation rather than an established finding.
But the trap is obvious: “Here are twenty ways to improve your paper” does not mean “this paper deserves to be published.”
Sometimes the most useful advice is that the central idea is insufficiently novel, the study cannot answer the question being asked, or the paper is simply not worth another month of work.
AI is exceptionally good at helping us improve things. It may be rather less good at telling us when to stop.
The bigger problem: AI can also be manipulated
A 2026 JAMA Network Open experiment inserted invisible white-on-white instructions into manuscripts directing three commercial LLMs to assess them favourably. Under neutral reviewing prompts, acceptance recommendations increased from 0% to nearly 100%.
The manipulation also reduced detection of genuine scientific flaws. Under stringent criteria, overall flaw detection fell from 18.9% to 8.5%, while detection of methodological flaws fell from 56.3% to 25.6%.
So what should the future of peer review look like?
AI should handle the tick-boxing. Does the abstract agree with the Results? Are the numbers internally consistent? Are all figures cited? Do Methods and Results contradict one another? Are outcomes reported consistently? Are references obviously missing? Has the relevant reporting guideline been followed? AI can also flag statistical assumptions that deserve examination, although flagging a possible problem is not the same as validating the statistics.
Confidentiality still matters
Submitted manuscripts are confidential. Nature Portfolio currently instructs reviewers not to upload manuscripts into generative-AI tools and requires disclosure when AI has supported evaluation of a manuscript’s claims.
So AI cannot simply be bolted onto peer review by asking reviewers to upload papers into general-purpose external LLMs. Any useful AI-assisted workflow must respect the same confidentiality obligations as human reviewers.
If such checks can be performed reliably within appropriately governed systems, using expert reviewer time to do them manually makes little sense.
The reviewer should concentrate on the questions for which expertise was sought: Does the science make sense? Is the finding genuinely novel or meaningful? Do the data justify the conclusions? Is the requested additional experiment actually necessary?
The distinction is familiar in dermatopathology. An immunostain may strongly support a diagnosis. But if the morphology and clinical picture do not fit, the stain does not get to make the diagnosis.
AI can provide evidence, identify inconsistencies and raise questions. The reviewer decides what those findings mean and remains responsible for every criticism sent to the author.
Editors remain essential too.
They must decide whether the reviewer understood the paper, whether the criticisms matter, whether revision requests are proportionate and whether another revision is justified at all.
I would much rather receive one page from an expert who understood my paper than three pages of immaculate commentary discussing everything except what is actually wrong with it.
The Future
AI will not make peer review obsolete. It should make the mechanical part of peer review obsolete.
AI-assisted peer review is already here, whether journals officially admit it or not. Rather than pretend otherwise, journals should legitimise it, regulate it, and build it properly into the review process.
Let AI do the checking.
Let reviewers do the thinking.
Let editors make the decision.
This, my friends, will be the future of peer review…..
References
Frontiers Media. Unlocking AI’s untapped potential: responsible innovation in research and publishing. Frontiers Media; 2025. Global survey of 1,645 active researchers. The report found that 53% of peer reviewers surveyed were using AI tools. Frontiers
Frontiers reportKeuper J. Prompt Injection Attacks on LLM Generated Reviews of Scientific Publications. In: Pattern Recognition: 28th International Conference, ICPR 2026. Lecture Notes in Computer Science, vol 16812. Springer, Cham. pp. 122–134. doi: 10.1007/978-3-032-31335-5_9. The study generated 1,000 reviews of ICLR 2024 papers and found acceptance recommendations above 95% for many LLMs even under neutral conditions. Springer
Springer publicationRusso Latona G, Horta Ribeiro M, Davidson TR, Veselovsky V, West R. The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates. Proc ACM Hum Comput Interact. 2025;9(7). doi: 10.1145/3757667. This is the source for the estimated 15.8% prevalence of AI-assisted ICLR 2024 reviews and the 4.9-percentage-point increase in acceptance among papers near the threshold. Princeton University
Published article/DOIChoi B, Jun TJ, Sung JW, et al. Invisible Text Injection and Peer Review by AI Models. JAMA Netw Open. 2026;9(1):e2552099. doi: 10.1001/jamanetworkopen.2025.52099. This supports your statements about invisible prompt injection increasing acceptance recommendations from 0% to nearly 100% and reducing detection of genuine flaws. JAMA Network
JAMA Network Open articleNature Portfolio. Nature Portfolio Journals Guide to Reviewers. Nature Portfolio. Accessed October 2, 2026. The guidance states that reviewers should not upload manuscripts into generative-AI tools and should disclose AI assistance when it supports evaluation of manuscript claims. Nature
Nature Portfolio reviewer guidance






