LINKEDIN HOOK
A language model handed me six perfectly formatted references. Two were fabricated. Here is what that means for anyone writing a manuscript — and the verification protocol I have never skipped.
A reference that did not exist
Some time ago I asked a language model for the literature on a question I already knew well. It returned six references. Correct format, plausible journals, plausible author groups, volume and page numbers that looked exactly like volume and page numbers. Two of them did not exist. Not wrong — absent. No such paper had ever been written.
I knew the field, so I caught it in under a minute. A resident writing her first review would not have caught it at all. She would have pasted it into her draft, and a fabricated citation would have entered the literature wearing a coat and tie.
This is the single most important thing to understand about using these tools for academic work, and it is why I am starting the useful part of this series with the failure rather than the benefit.
Why it happens
It is not a bug and it will not be patched away. A language model produces text that fits the pattern of what came before. A reference list has an extremely regular pattern. Authors, then title, then journal, then year, then volume, then pages, then a DOI that looks like every other DOI. The model can generate that shape perfectly without any of the underlying records existing, because generating the shape is the entire task it was trained on.
It follows that the model is at its most dangerous precisely where academic work is most rule-bound. The bibliography. The methods section. The statistical phrasing. Anywhere the surface form is highly regular, fluency and accuracy come apart completely.
So the rule I work by is simple and has no exceptions. The model never supplies a reference. It may help me find one. I verify it, or it does not go in.
The verification protocol below is the part I have never once skipped. Paid subscribers read on.


