If you are an academic physician, try this experiment.
Upload a manuscript to a capable AI model and type:
“Review this paper for me.”
You may get an impressive response.
You may also get three pages of minor complaints, stylistic suggestions, speculative concerns, and methodological objections that sound sophisticated but do not materially affect the paper.
The problem is not necessarily the AI.
The problem may be your prompt.
We would never invite a human reviewer to evaluate a manuscript without telling them what we need. Is this an editorial triage? A statistical review? A methodological review? A search for fatal flaws? A review for publication? Are we interested in grammar, or only problems that could change the conclusions?
Yet we routinely give AI four words:
“Review this for me.”
Then we judge the model by what comes back.
That is not AI competence.
Tell AI what matters
For most manuscript reviews, I do not want twenty minor criticisms.
I want to know:
Is there anything seriously wrong with this paper?
I want the model to distinguish between a problem that threatens the validity of the conclusions and something that merely could have been done differently.
So instead of:
Review this paper for me.
try this:
Act as a rigorous peer reviewer for a high-quality medical journal. Review the attached manuscript for MAJOR problems only.
Focus on errors or limitations that could materially affect the validity, interpretation, reproducibility, or clinical implications of the study.
Specifically examine:
Study design and whether it can answer the stated research question.
Selection bias, misclassification, confounding, and other important sources of bias.
Statistical methods and whether the analyses support the conclusions.
Discrepancies between the data presented and the authors’ claims.
Unsupported causal language or conclusions that go beyond the data.
Missing analyses or information that could materially change interpretation.
Internal inconsistencies among the abstract, text, tables, figures, and conclusions.
Clinically important limitations that the authors have overlooked or understated.
Do not list minor stylistic issues, preferences, or changes that would not materially affect the paper.
For every major concern:
identify exactly where the problem occurs;
explain why it matters;
distinguish a demonstrated error from a possible concern;
state what evidence in the manuscript supports your criticism; and
explain what would be required to correct or resolve it.
Do not invent missing information. If something cannot be determined from the manuscript, say so explicitly.
End with:
Overall assessment: Are there major problems that materially weaken the paper?
Recommendation: Acceptable as written, minor revision, major revision, or not suitable for publication, with a brief explanation.
That is a very different assignment.
But a better prompt does not make AI right
This is the part of AI competence that is often forgotten.
A sophisticated prompt can produce a sophisticated-looking error.
AI can misunderstand the study design. It can criticize an analysis that is actually appropriate. It can overlook an important confounder. It can claim that information is missing when it is sitting in Table 3.
And, unless constrained and checked, it can fabricate references or confidently attribute claims to papers that do not support them.
So the output is not the peer review.
It is a candidate analysis for the reviewer to interrogate.
For every important criticism, ask:
Where exactly is the evidence for this?
Then go back to the manuscript.
If the criticism depends on an external reference, open the original reference.
If the model says the statistics are wrong, verify the statistical argument.
If it says the authors have contradicted themselves, compare the passages yourself.
The clinician or scientist remains responsible for the final judgment.
This is Clinical AI Competence
Prompt engineering is useful.
But clinical AI competence is larger than prompt engineering.
It means knowing what to ask, what to trust, what to verify, and when not to accept the machine’s answer.
That distinction will become increasingly important as AI gets better.
The danger is not merely that AI will produce obviously bad answers.
The more interesting problem is the opposite:
AI will increasingly produce answers that are so persuasive that we may forget to check whether they are true.
That is why physicians and scientists need more than access to AI.
They need competence in using it.
And sometimes that competence begins with something as simple as replacing:
“Review this paper for me.”
with a much better question.
Clinical AI Competence #1
This is the first in a continuing ObGyn Intelligence series on using AI effectively and safely in clinical medicine, research, education, and academic work.
Subscribe to ObGyn Intelligence for the next Clinical AI Competence lesson.


