The important comparison is not AI versus physicians. It is physicians writing alone versus expert physicians writing with AI.
I used to think the debate over AI and writing could be reduced to a simple question:
Can AI write as well as humans?
I no longer think that is the interesting question.
For physicians, scientists, and medical educators, the question that matters is:
Can an expert clinician write better with AI than without it?
Increasingly, I think the answer may be yes.
Not because AI knows medicine better than physicians. It does not.
Not because AI should be trusted to determine what is scientifically true. It should not.
And certainly not because physicians should surrender authorship, judgment, or responsibility to a machine.
The reason is much simpler.
Writing and knowing are different cognitive tasks.
A physician may know exactly what needs to be said and still write it imprecisely. We bury the main point. We repeat ourselves. We use ambiguous terminology. We mix observation with inference. We overstate causality. We forget to define the denominator. We write a conclusion stronger than the data justify.
Expertise does not automatically produce excellent prose.
AI is unusually good at some of those mechanical and structural tasks.
The evidence is becoming interesting
A recent study compared two scientific manuscripts constructed from the same title, abstract, and tables. One was written by an experienced physician and the other by GPT-4o.
The AI manuscript scored 9.0 versus 7.2 for clarity and readability.
But there was an equally important finding.
The physician-written manuscript was substantially better in technical accuracy, 9.3 versus 6.3, and also better in depth. PubMed
That result does not tell me that AI should write medical papers.
It tells me something much more useful:
AI and physicians appear to have complementary weaknesses. And strengths.
AI can make prose orderly, explicit, concise, and readable.
The expert clinician supplies what AI cannot safely manufacture: scientific judgment, clinical context, interpretation, skepticism, and accountability.
Put the two together and we may have a better writing system than either alone.
We see the same signal in clinical communication
In a widely discussed JAMA Internal Medicine study, health professionals preferred ChatGPT responses over physicians’ answers to patient questions in 78.6% of evaluations. Responses from the chatbot were also rated substantially higher for quality and empathy. JAMA Network
That study was not a trial of clinical care, and the AI answers were much longer. It should not be overinterpreted.
But subsequent work has produced similar signals.
At NYU Langone, physicians evaluating AI-generated patient-message drafts rated their communication style better than human-generated messages, although information quality was similar. AI messages were also more complex linguistically, demonstrating that AI itself requires editing and supervision. JAMA Network
The lesson is not “AI wins.”
The lesson is that AI can systematically improve certain dimensions of communication.
And now we can see the structural difference
A fascinating new preprint examined 2,250 human-written commercial articles and 11,250 matched articles generated by five frontier AI models.
AI writing had a remarkably recognizable structure.
It tended to state its thesis early, announce where the argument was going, organize information into explicit stages, explain systematically, and finish by synthesizing the argument.
Human writing was much more structurally unusual and heterogeneous. 2609.15369 (1)
That was presented partly as a way of identifying AI-generated writing.
But I see another implication.
Many of the characteristics that make AI writing detectable are precisely the characteristics we teach authors to use:
State the question (it’s in the prompting!).
Define the terms.
Organize the argument.
Distinguish evidence from interpretation.
Tell the reader what follows.
Finish with a conclusion that actually follows from the evidence.
Being recognizably AI-like is not necessarily the same thing as writing badly.
Medicine is a special case
I am particularly interested in this because medical writing is not primarily creative writing.
Its purpose is usually not stylistic originality.
Its purpose is precision.
A medical manuscript should make it difficult for the reader to misunderstand:
what population was studied,
what intervention or exposure occurred,
what comparison was made,
what outcome was measured,
what the numbers actually show,
what remains uncertain,
and what conclusions the evidence does and does not justify.
That makes medicine unusually well suited to AI-assisted writing.
AI can ask:
Did you define this term?
Is that a causal claim from observational data?
Does the conclusion exceed the results?
Is the denominator clear?
Are the numbers in the abstract consistent with Table 2?
Does “significant” mean statistically significant or clinically important?
Does this recommendation actually follow from the cited evidence?
Those are not trivial editorial improvements.
They can be patient-safety improvements.
But there is a non-negotiable condition
AI-generated medical writing cannot be trusted simply because it sounds precise.
LLMs can fabricate citations, introduce unsupported statements, erase important nuance, and confidently transform uncertainty into apparent fact. Comparative studies continue to demonstrate citation and factual errors. ScienceDirect
So the responsible model is not:
AI writes. Physician approves.
It is:
Physician thinks. AI challenges, structures, checks, and edits. Physician verifies.
That final word matters.
Verifies.
I increasingly use AI not simply to generate sentences but to attack my own writing.
Find the unsupported claim.
Find the denominator error.
Find the contradiction between text and table.
Find the sentence that implies causality without evidence.
Find the reference that does not support the claim attached to it.
Find what a skeptical reviewer will find before the reviewer finds it.
That is a very different conception of AI.
It is not a ghostwriter.
It is an intellectual instrument.
The future medical writer may therefore be neither human nor AI
It will be the clinician who knows how to combine both.
The physician contributes expertise, judgment, experience, skepticism, ethics, and responsibility.
AI contributes extraordinary capacity for structure, comparison, consistency checking, linguistic revision, and tireless interrogation of a document.
Neither is sufficient alone.
But used properly, the combination may produce medical writing that is clearer, more disciplined, and more precise than what most of us produce unaided.
That is why I believe clinical AI competence is becoming part of professional competence.
We should stop asking whether AI can write like doctors.
The more consequential question is:
Why would a doctor who can use AI competently choose to write without it?
References
Ayers JW, Poliak A, Dredze M, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med. 2023;183:589-596. JAMA Network
Can artificial intelligence write science? A comparative analysis of human-written and artificial intelligence-generated scientific writings. 2025. The study found greater AI clarity/readability but lower technical accuracy and depth. PubMed
Large language model-based responses to patients’ in-basket messages. JAMA Netw Open. 2024. JAMA Network
Madler J. SlopShape: Identifying AI-Generated Commercial Web Content. arXiv preprint, 2026. 2609.15369 (1)
I particularly like the final question. It moves the discussion away from the rather sterile “human versus AI” competition and toward the professional obligation you have been developing: competent clinicians should know how to use AI while retaining responsibility for every claim.


