Last week I read a new consult series from a major professional Ob society. It carries the endorsement of a second society. Thirty-five pages. Two hundred eighty-four references. The kind of document that tells doctors what to do on Monday morning.
I did not set out to audit it.
Something in one of the radiation tables refused to sit still. Two tables sat side by side, under identical headers, and disagreed with each other. When I checked them row by row, they were offset by exactly two weeks. One counted from conception. The other counted from the last menstrual period. Nobody had noticed.
So I built a prompt and pointed it at the whole document. It came back with nineteen items.
A steroid dose conversion that used two different ratios in the same sentence. They cannot both be right, and neither one is. A recommended waiting time between chemotherapy and delivery that matched exactly the waiting time its own cited study had identified as the worst one. A cancer stage that cannot exist under the current staging system. A study cited as measuring time from one event when it actually measured time from a different one, in a direction that reverses the advice a woman would be given.
Not one of these was a mistake about obstetrics.
Every single one sat on a border with another field. Radiation physics. Blood platelet thresholds. Drug potency. Cancer staging. Lymphoma staging.
That is where documents written by one specialty about five others break.
Why I am writing about prompting instead
I could not have done this three years ago. I could not have done it last year, not this well. What changed is not that I got better at reading. What changed is the machines, and almost everything doctors have been taught about how to talk to them is now out of date.
The old advice was to learn tricks.
Give the model a persona.
Show it examples.
Tell it to think step by step.
In 2023 that advice was correct.
A Microsoft team showed that a stack of prompting tricks, which they called Medprompt, pushed a general model past specialist medical models on nine medical benchmarks and above 90 percent on the medical licensing question set for the first time. Note that this work is a preprint and has not been peer reviewed. (1)
Then the reasoning models arrived, and the same team found the opposite. With the newer models, adding examples to the prompt made performance worse, not better. Their words: in-context learning may no longer be an effective steering approach. Also a preprint. (2)
Peer-reviewed work now says the same thing. A 2025 study compared several chain-of-thought prompting methods across medical question sets and found no significant difference between them. The authors concluded that complex prompting techniques do not significantly improve performance compared with simpler approaches. (3) Another group tested prompting strategies across three models on clinical decision tasks and found the effect went in both directions. It helped the weakest task and was counterproductive for others. That one is also a preprint. (4)
So the tricks are fading. Three things replaced them, and they matter more than any trick ever did.
What actually works now
The first is telling the model what job it has and what the answer should look like. Not clever wording. Plain specification. The reasoning models do their own thinking; what they need from you is the task, the constraints, and the format.
The second is telling it to prioritize safety, out loud, in the prompt. This is not decoration. In a 2025 study of three reasoning-capable models across clinical scenarios graded by six expert clinicians, critical safety problems appeared in 12 of every 100 responses overall. A safety-first instruction cut those from about 16 in 100 down to about 9 in 100. That is a 45 percent reduction from one sentence. The same study found that even the best prompting left the models poor at communication and empathy, at roughly half of the maximum score, and concluded that current prompt engineering gives only marginal improvement, not enough for reliable clinical use. (5)
The third is verification. Tell the model to check its own claims against the original source, to quote what it found, and to mark plainly what it could not verify. This is the whole reason my audit prompt works. It is not clever. It forces the model to label every finding as proven from the document, proven against a retrieved source, or unverified. The unverified ones I take to the authors as questions, not accusations.
The fourth (optional) one is to enter the finding into a different AI model and ask it to verify the results of the first one
One practical note before you paste it. The free tier of most of these services gives you a smaller model, less thinking time, and a shorter memory, and none of that is enough for a thirty-five page document with 284 references. The gap is measurable: on board-style medical questions this year, models that reason step by step scored 81.5 percent against 75.8 percent for those that do not, and in an oncology text-mining study the same model's accuracy rose from 0.84 to 0.94 simply by being allowed to think longer. If you are going to check a document that tells doctors what to do on Monday, pay for the better model.
The failure nobody warns you about
Here is the most important study of the last year, and it is not about prompting at all.
Researchers gave 1,298 people ten medical scenarios and asked them to work out the likely condition and what to do. Some got help from a large language model. Some used whatever source they liked. Tested on its own, the model identified the right condition 94.9 percent of the time. The people using that same model identified it fewer than 34.5 percent of the time, which was no better than the people with no model at all. (6)
Read that again. The machine knew. The human did not get it out.
The authors located the failure precisely. Their words: the transmission of information between the model and the user is a particular point of failure. People gave the model incomplete information, and the model produced the right answer but did not get it across. They also noted that standard medical benchmarks did not predict this at all. (6)
There is a second version of this problem aimed straight at us. A 2026 evaluation across more than 5,500 medical questions found that models readily abandoned correct reasoning when they were given misleading hints attributed to an expert physician. (7) So the model can be talked out of a right answer by a confident doctor. If you use one of these tools and push back on it because it disagrees with you, understand what you may be doing.
What this means for clinicians and patients
Do not treat the tool as the doctor. Treat it as the reader who never gets tired at reference number 140. That is a real job and we are demonstrably worse at it.
Be honest that it is not uniformly better. When one group asked models to agree or disagree with an orthopedic guideline, the best prompt reached only 62.9 percent agreement overall. (8) When another group ran 300 manuscripts past both models and 324 ophthalmologists, the humans rejected 73 of every 100 papers and the models rejected 2, repeating the same three generic complaints and omitting line numbers and references. Those authors concluded the tools should not be used for manuscript revision in their current state. (9) Both findings are real and I am not going to hide them.
But in the other direction, a blinded cardiology study this year found model-generated reviews scored higher than human reviews in five of seven quality domains, agreed with the final editorial decision about as often, and took 2 to 6 minutes against a median human turnaround of 17 days. (10) A doctor who dismisses all of this is not being careful. He is being comfortable.
For a patient, the practical message is short. If you ask an AI about your symptoms, the answer you receive is probably worse than the answer the machine actually had, because the handoff is where it breaks. Tell it everything. Ask it what it would need to know to be more certain. Then take what it says to a human.
Conclusion
The prompting advice circulating right now is a year or two behind the models it describes. Personas and clever phrasing are not where the value is anymore. Specification, an explicit safety instruction, and forced verification are, and the largest single source of error is not the machine at all. It is the space between the machine and the person using it.
For our profession the conclusion is sharper. A guideline endorsed by two societies, reviewed at multiple levels, went to press with a dose conversion that fails its own arithmetic and a cancer stage that cannot exist. Nobody checked the numbers against the sources, because by hand nobody can. That is no longer an acceptable reason. Every professional society document should be machine-checked for arithmetic, units, and citation fidelity before publication. Not machine-written. Machine-checked. The two are not the same thing and the distinction is the entire argument.
I am not asking anyone to trust the output. I checked every finding myself, and where a paper sat behind a paywall I could not open, I marked the item unverified instead of claiming it. That is the division of labour that works.



