Imagine reviewing a manuscript in your specialty.
The question matters. The methods appear reasonable. You examine the analysis, challenge the conclusions, and recommend revisions. The authors respond. The paper is accepted.
Buried in its bibliography is a study that does not exist.
The citation looks entirely ordinary. Recognizable journal. Plausible title. Familiar author names. Nothing about its appearance tells you that the evidence it supposedly represents is missing.
Your review may have improved the manuscript substantially. Yet something more basic escaped scrutiny: whether part of its supporting literature was real.
That possibility is now the subject of a troubling research report.
And it leads me to a position that journals need to take seriously: AI assistance must become part of the infrastructure of peer review.
This is my proposed standard, not a claim that every available AI tool has been proved effective.
In their 2026 arXiv preprint, Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences, first author Mark Russinovich and colleagues audited accepted papers from major AI and security conferences. Their definition covered nonexistent works and substantial author mismatches, excluding routine bibliographic discrepancies. Approximately one in twenty NeurIPS and USENIX Security papers in 2025 contained at least two likely hallucinated academic references. Reference-level rates were generally below 1%.1
The authors interpreted this as evidence that peer review alone does not reliably enforce citation integrity. Their verification pipeline combined bibliographic databases with AI-assisted web searching.1
The limitations matter. This is a preprint, not an established estimate for medical journals. The audit does not establish that AI generated each problematic reference, and citation identity checks do not establish whether a real source supports a claim.1
My conclusion is that journals should make systematic verification an explicit part of review.
Expertise cannot substitute for retrieval.
An experienced reviewer may recognize an invented landmark trial immediately. But a plausible reference to an unfamiliar paper presents a different problem.
You cannot establish that a paper exists by knowing a great deal about its subject. You must retrieve it.
You cannot establish that it supports a sentence by recognizing the journal. You must examine what it actually reports.
You cannot establish that its findings apply to the population under discussion by reading its title. You must compare the study with the claim.
A reviewer can perform these checks for individual references. The unrealistic expectation is that every reviewer will perform them exhaustively, for every submission, while also evaluating design, statistics, interpretation, originality, and clinical relevance.
No individual reviewer can serve as an exhaustive verification system for the scientific literature at publishing scale.
Telling reviewers to “be more careful” does not solve that design problem.
Consider a hypothetical manuscript with 80 references. Allow just three minutes to locate and check each one. That is four hours before assessing whether the sources support the statements attached to them. The arithmetic is illustrative, but the workload is real in principle: every additional layer of verification consumes time.
The response should be to build a process that makes those checks routine.
The objection that “AI hallucinates” strengthens the case for verification.
It is a valid objection to treating a chatbot’s answer as evidence.
Ask an AI system whether a citation is real, accept its confident answer, and you may simply add another unsupported assertion to the chain.
A useful verification system must show its work in an inspectable form: the retrieved record, the matching publication, the relevant passage, and any unresolved discrepancy.
The distinction matters. A model’s recollection is not a bibliographic record. Agreement between two models is not independent confirmation. A functioning link does not establish that the linked paper supports the manuscript’s claim.
AI assistance should help retrieve, compare, and flag. Its conclusions must remain open to inspection.
Straightforward checks should use conventional software and authoritative databases wherever possible. AI adds a potential layer of assistance when references are incomplete, wording differs, or a claim needs comparison with source text. Those more interpretive functions require separate validation.
The goal is a documented chain from assertion to evidence.
A real reference can still support a false narrative.
Imagine a manuscript stating that an intervention “prevents complications.”
The reference exists. The authors and journal are correct. The DOI works.
But the cited study measured a laboratory marker rather than complications. Or it enrolled a different population. Or its observational design supports an association while the manuscript asserts a causal effect.
These are hypothetical examples of the questions a review process should ask. Passing a citation-existence check would leave every one of them unresolved.
That is why I would require two distinct layers: bibliographic verification for every reference, followed by source comparison for the claims that carry the paper’s argument.
AI can assist with that comparison by placing the manuscript’s statement beside the relevant source passage and identifying possible discrepancies. The reviewer then assesses whether the discrepancy matters.
A system that presents the evidence for scrutiny is much more useful than one that merely produces another polished review.
Journals should provide the tools and own the process.
My proposed standard would require journals to:
Check every reference against retrievable bibliographic records.
Provide reviewers with an AI-assisted evidence audit for the manuscript’s central claims.
Distinguish confirmed errors from unresolved searches and interpretive concerns.
Require human assessment before an automated flag influences an editorial decision.
Validate the system and monitor both missed errors and false accusations.
An unsuccessful search should be labeled “unresolved.” Calling it “fabricated” requires stronger evidence. Authors must have a fair opportunity to supply the source or correct the record.
Confidentiality also belongs in the design. ICMJE requires reviewers to follow journal AI policies or obtain permission, preserve manuscript confidentiality, disclose AI use, and ensure that the resulting content is appropriate and valid.2
Journals should therefore provide approved, secure systems. Reviewers should not have to improvise access to the tools the journal expects them to use.
The responsibility remains human. The assistance should be systematic.
Authors remain responsible for their references and claims. Reviewers remain responsible for their judgments. Editors remain responsible for publication decisions.
Adding AI does not transfer any of those obligations.
It can, however, give those people a more complete set of questions to investigate and a clearer record of what has actually been checked.
The conference audit does not prove that AI-assisted review improves clinical outcomes or catches every false claim. Those outcomes need evaluation. It does expose a verification gap that deserves a concrete response.
My position is that journals must build and evaluate AI-assisted verification as a standard part of review, with human oversight and traceable evidence. They should define its tasks narrowly enough to test its performance and expand its role only when the results justify doing so.
“Peer reviewed” should never be mistaken for “every claim verified.”
But we should be working to close that gap.
AI hallucinations make the need harder to ignore. A fabricated reference can look perfectly scholarly. Human expertise cannot make an unread source real.
Reviewers need tools that help them discover what deserves closer inspection.
Journals must provide them.
References
Russinovich M, Siva Kumar RS, Salem A. Phantom references: hallucinated citations that survive peer review at top-tier conferences. arXiv [Preprint]. 2026 Jul 1. arXiv:2607.00738. Full text.
International Committee of Medical Journal Editors. Use of AI by reviewers. In: Recommendations for the conduct, reporting, editing, and publication of scholarly work in medical journals [Internet]. [cited 2026 Sep 8]. Recommendations.


