Around 1970, obstetrics embraced continuous electronic fetal monitoring with an intuitively compelling proposition. If clinicians could watch the fetal heart rate continuously, rather than intermittently, they should recognize fetal hypoxia earlier, intervene sooner, prevent neurologic injury, and ultimately reduce cerebral palsy.
The physiology made sense. The technology produced objective-looking data. The tracing could be preserved in the medical record. Continuous monitoring also appeared more sophisticated than intermittent auscultation.
Adoption therefore moved faster than definitive evidence of improved clinical outcomes.
Half a century later, continuous electronic fetal monitoring is used in more than 85% of American labors. Yet the outcome it was expected most importantly to prevent, cerebral palsy, did not undergo the dramatic decline anticipated with widespread monitoring. Meanwhile, operative delivery increased substantially over the same era, although fetal monitoring was obviously not the only reason.
The randomized evidence remains sobering. Compared with intermittent auscultation, the Cochrane review found that continuous cardiotocography reduced neonatal seizures, relative risk 0.50, but demonstrated no reduction in perinatal death or cerebral palsy. Cesarean delivery increased by about 60%, and operative vaginal delivery also increased. Much of this evidence is old and of limited quality, which itself is part of the problem. A technology became nearly universal without the kind of contemporary outcome evidence we would now expect before introducing many far less consequential interventions.
The consequences are visible today. Provisional CDC data for 2025 show that the overall US cesarean delivery rate increased again, to 32.5%. Among nulliparous, term, singleton, vertex births, the rate reached 26.9%, its highest level since 2012. Continuous fetal monitoring cannot be assigned responsibility for those rates, but it remains central to the chain of events leading to many intrapartum cesareans. (CDC)
This history matters now because obstetrics is standing at the beginning of another technological transformation.
This time the machine is artificial intelligence.
And unlike electronic fetal monitoring, AI will not take 20 years to become ubiquitous.
In the American Medical Association’s 2026 survey of 1,692 physicians, 81% reported either using AI professionally or awareness of AI use in their practice, compared with 38% reporting use in 2023. The average number of AI use cases had more than doubled. Research summarization and clinical documentation were already among the most common applications. AI is not approaching medicine. It is already inside it. (American Medical Association)
The relevant question is therefore no longer whether obstetricians will use AI.
It is what evidence we should require before AI influences decisions that change what happens to women and fetuses.
Four different technologies hiding under one name
“Artificial intelligence in obstetrics and gynecology” is almost meaningless as a category.
Software that assists fetal ultrasound interpretation is not clinically equivalent to software predicting fetal acidemia from a cardiotocograph. Neither is equivalent to an algorithm ranking embryos, an ambient system writing a prenatal visit note, or a large language model advising a physician about management.
These technologies operate at different points in the causal pathway to patient outcomes. They should therefore earn different levels of trust.
The mistake would be to treat “AI” itself as the intervention.
Ultrasound may be one of the strongest candidates
Fetal ultrasound is particularly attractive for AI because it is already an image-recognition discipline, and diagnostic performance depends considerably on operator expertise, image acquisition, recognition of uncommon findings, and human attention.
The evidence is becoming interesting.
An FDA-cleared fetal ultrasound system designed to identify abnormal findings has now been evaluated across multiple settings. In a 2026 Obstetrics & Gynecology study involving 6,452 images from approximately 1,000 pregnancies at 75 sites in five countries, the system achieved mean sensitivity of 93.2% and specificity of 90.8% across eight predefined abnormal fetal findings. Performance was relatively stable across imaging equipment, geography, maternal BMI, and gestational age. These are encouraging results, although several investigators had financial relationships with the manufacturer. (PubMed)
More importantly, a separate multi-reader study published in npj Digital Medicine in July 2026 asked the more clinically relevant question: does AI make clinicians better?
That is a substantial step forward. The important unit of analysis is not merely whether an algorithm can identify an abnormal image. The clinically meaningful question is what happens when a clinician uses it.
But even a successful reader study remains an intermediate endpoint.
Higher sensitivity is not yet a healthier fetus.
The next questions are whether AI identifies anomalies that would otherwise have been missed, whether diagnoses are made early enough to change counseling or management, whether false-positive referrals increase, whether expert review corrects inappropriate AI alerts, and ultimately whether maternal or neonatal outcomes improve.
That hierarchy of evidence matters.
Algorithm performance → clinician performance → management change → patient outcome.
We should stop pretending those are interchangeable.
Fetal monitoring should make us much more cautious
Few areas illustrate the distinction better than cardiotocography.
In 2017, the INFANT Collaborative Group performed exactly the sort of trial obstetrics should remember. Investigators randomized 47,062 women undergoing continuous fetal monitoring at 24 maternity units in the United Kingdom and Ireland to computerized interpretation of the tracing with decision support or standard monitoring without that decision support.
The result was unequivocally negative.
Poor neonatal outcome occurred in 0.7% of babies in each group: 172 versus 171 events. The adjusted risk ratio was 1.01, with a 95% confidence interval of 0.82 to 1.25. At two years, there was no significant difference in developmental outcomes.
The computerized system interpreted information. It generated alerts. Clinicians received assistance.
Babies did not do better. (PubMed)
The INFANT trial should be required reading for anyone proposing AI-assisted fetal monitoring.
Yet the technology continues to improve, and newer data are revealing for another reason.
A randomized multi-reader study published in 2026 asked 211 obstetricians, residents, and midwives from 23 countries to assess fetal heart rate tracings with or without computerized assistance. The test set was intentionally enriched: half of the 100 cases had an umbilical cord pH below 7.15.
Without computerized assistance, clinicians correctly classified the cases only 54.0% of the time.
With assistance, accuracy increased to 61.4%.
Sensitivity improved from 49.3% to 61.7%, while specificity did not significantly decline. (Nature)
That improvement is real.
But the more disturbing finding may be the baseline.
Experienced clinicians evaluating a test that has been central to modern obstetrics for half a century performed only modestly better than chance in this deliberately balanced dataset.
AI exposed not only what the machine can do, but what humans cannot reliably do.
And there is an additional problem. In actual labor populations, severe neonatal acidemia is uncommon. A dataset in which 50% of babies are acidemic is useful for testing discrimination, but it cannot tell us what positive predictive value, alert burden, intervention rate, or false-positive consequences will look like on a real labor floor.
That requires prospective clinical trials.
We should demand them.
Embryo selection gives us another warning
Reproductive medicine offers a different example of how seductive AI performance can become.
Embryo selection appears almost designed for deep learning. Embryologists evaluate complex visual information. Embryo morphology contains thousands of subtle features. Computers should theoretically identify patterns invisible to human observers.
The hypothesis is reasonable.
But reasonable hypotheses still require trials.
A double-blind randomized noninferiority trial published in Nature Medicine compared deep-learning embryo selection with conventional morphology-based selection. Clinical pregnancy occurred in 46.5% of women assigned to deep-learning selection and 48.2% with standard embryologist assessment. The trial did not establish the prespecified noninferiority of the AI approach. Other pregnancy outcomes were also similar. (Nature)
That does not mean AI will never improve embryo selection.
It means it had not done so in that trial.
Those are very different statements.
An algorithm can predict implantation in a retrospective dataset, generate impressive receiver operating characteristic curves, rank embryos reproducibly, and still fail to improve the outcome women actually came to an IVF clinic to achieve.
Prediction is not treatment.
Documentation may be the quiet success story
Interestingly, one of medicine’s least glamorous AI applications may currently have some of its most reproducible practical evidence.
Ambient AI scribes listen to clinical encounters and prepare draft documentation.
Here the causal pathway is short.
If the intended outcome is reduced documentation burden, we can measure documentation burden.
Pragmatic randomized trials have begun to show reductions in time spent writing notes, improvements in workflow satisfaction, and reductions in several measures related to professional exhaustion. A 2026 randomized crossover trial of 160 outpatient clinicians found improvements in satisfaction and burnout measures with two ambient systems, although effects on some objective efficiency metrics were modest. A separate pragmatic randomized trial of 238 physicians found a significant reduction in time-in-note with one of two systems and improvements in task load and work exhaustion among AI-scribe users. (PubMed)
The effects are not revolutionary.
That may be precisely why they are credible.
AI does not have to diagnose fetal compromise or discover a congenital anomaly to be useful. Removing repetitive administrative work may produce more reliable value than attempting to replace high-stakes clinical judgment.
The least dramatic application may therefore be among the most defensible.
Breast screening shows what serious evaluation looks like
If obstetrics wants a model for evaluating clinical AI, it should look outside obstetrics.
The Mammography Screening with Artificial Intelligence, or MASAI, trial randomized 105,934 women in the Swedish national breast-screening program to AI-supported screening or conventional double reading.
The initial results were impressive but, importantly, investigators kept following the participants rather than declaring victory after an improvement in an AUC.
AI-supported screening detected 29% more cancers, 6.4 versus 5.0 cancers per 1,000 women screened, without a significant increase in false-positive results. Reading workload fell by 44.2%. The additional cancers detected were predominantly small, lymph-node-negative invasive cancers. (DOI)
Then came the harder endpoint.
With follow-up, the interval cancer rate was 1.55 per 1,000 in the AI group compared with 1.76 per 1,000 with standard screening. Sensitivity increased from 73.8% to 80.5%, while specificity remained 98.5% in both groups. There were also descriptively fewer invasive, larger, and biologically unfavorable interval cancers in the AI-supported group. (PubMed)
That is what serious evaluation looks like.
Large population.
Randomization.
A real clinical workflow.
Real clinicians.
False positives measured.
Workload measured.
Harder downstream outcomes measured.
Follow-up measured in years rather than minutes.
That is a much higher standard than demonstrating that an algorithm can outperform a physician on a selected image bank.
It should become our standard as well.
Three questions every AI study should have to answer
Every clinical AI claim in women’s health should face three questions:
Compared with whom?
Tested in what population?
Measured by what outcome?
Those questions sound elementary, but much of the AI literature becomes considerably less impressive once they are asked.
An AUC is not a maternal or neonatal outcome.
Accuracy is not an outcome.
Agreement with experts is not an outcome.
A reduction in documentation time can be an outcome if documentation burden is what the intervention is designed to address.
An anomaly detected that would otherwise have been missed is closer to a clinical outcome.
A changed delivery plan that prevents neonatal morbidity is better.
A severe neurologic injury prevented is better still.
A live birth after IVF is more meaningful than embryo-ranking accuracy.
An unnecessary cesarean avoided without additional neonatal harm is more meaningful than improved classification of fetal heart rate tracings.
The closer we move toward the patient, the harder the study becomes.
That is not a reason to avoid the study.
It is the reason to do it.
AI will also expose the weaknesses of human medicine
There is another lesson emerging from this evidence.
The comparison cannot always be AI versus perfection.
Human performance is frequently poor.
Fetal monitoring interpretation has substantial interobserver variability. Congenital anomalies are missed on routine ultrasound. Embryo grading varies between observers. Physicians forget information, overlook abnormal results, make arithmetic mistakes, misremember guidelines, and suffer fatigue and distraction.
The proper comparator for AI is therefore usually actual clinical practice, not an imaginary infallible physician.
That is an argument for AI, not against it.
But it creates a reciprocal obligation.
If an AI system demonstrably and reproducibly performs a clinically important task better than physicians and improves patient outcomes, refusing to use it eventually becomes as ethically important a question as adopting an inadequately validated system too early.
Professional responsibility applies in both directions.
We should not adopt technology merely because it is new.
We should also not reject validated technology merely because the cognition occurs in silicon rather than cortex.
What AI cannot repair
There is another danger in allowing enthusiasm for medical AI to become too expansive.
Algorithms operate inside health systems.
They do not replace them.
The March of Dimes classified 1,104 US counties, 35.1% of all counties, as maternity care deserts in its 2024 report. These counties had neither a birthing facility nor an obstetric clinician and were home to more than 2.3 million women of reproductive age. (March of Dimes)
AI can improve ultrasound interpretation.
It cannot reopen a closed labor and delivery unit.
It can identify hypertension in an electronic record.
It cannot manufacture a postpartum nurse.
It can recommend urgent evaluation.
It cannot create transportation for a woman living 70 miles from a hospital.
It may eventually allow a rural clinician to obtain expert-level decision support, improve tele-ultrasound, detect deterioration earlier, and extend scarce specialist expertise over enormous geographic distances. Those could be major contributions.
But a model cannot substitute for the infrastructure necessary to act on its recommendation.
An algorithm that correctly identifies a high-risk pregnancy in a community without accessible obstetric care has diagnosed a systems failure. It has not fixed one.
That distinction is especially important when AI is promoted as a solution to maternal mortality.
Some problems are informational.
Others are structural.
We should know which one we are trying to solve.
The standard I would use
I am strongly in favor of medical AI.
I use it every day.
My concern about obstetrics is not that we are moving too quickly toward AI. In many areas, we are moving far too slowly in developing the competence required to use it intelligently.
But enthusiasm cannot substitute for evidence.
The history of electronic fetal monitoring gives us an unusually relevant warning. Obstetrics adopted a pattern-recognition technology because the physiological argument was compelling. It became embedded in clinical practice, hospital policy, medicolegal expectations, training, documentation, and eventually the definition of ordinary obstetric care.
By the time the outcome evidence was mature, the technology was effectively irreversible.
We should not repeat that sequence with AI.
I would require several things.
First, the evidentiary threshold should increase with the clinical consequence of the algorithm. An AI scribe does not require the same evidence as software recommending operative delivery. A system that changes whether an obstetrician performs a cesarean deserves prospective evaluation of maternal and neonatal outcomes, not simply validation against expert interpretation.
Second, AI should be evaluated as part of the human-AI clinical system, not only as isolated software. The clinically relevant question is rarely whether AI beats a doctor. It is whether a doctor using AI produces better decisions than a doctor without it.
Third, performance must be tested in the population in which the system will actually be used. Enriched case-control datasets are appropriate for early validation but cannot establish clinical utility.
Fourth, consequential AI use requires transparency. When an algorithm materially influences diagnosis, prognosis, embryo selection, fetal surveillance, or a recommendation for intervention, patients should not unknowingly become subjects of algorithmically mediated care. Exactly how disclosure should occur will depend on the application, but informed clinical decision-making should not acquire a software exception.
Fifth, postdeployment surveillance matters. AI systems can change. Clinical populations change. Workflows adapt around technology. Automation bias develops. A model performing well at introduction is not entitled to permanent trust.
And finally, we should resist allowing AI to become the standard of care merely through ubiquity.
The appropriate sequence is:
Validation. Clinical evaluation. Competency. Implementation. Surveillance. Standard of care.
Not:
Availability. Adoption. Habit. Standard of care. Evidence later.
We already ran that experiment once.
It lasted fifty years.
Artificial intelligence may ultimately become one of the most important advances in obstetrics and gynecology. It may detect abnormalities humans miss, reduce variability, rescue clinicians from impossible information loads, make expertise available where specialists do not exist, and perhaps eventually make some forms of clinical cognition demonstrably safer than unaided human judgment.
I expect much of that to happen.
But the ethical obligation is not to prove that AI is impressive.
It is to prove that patients are better off because we used it.
Electronic fetal monitoring taught obstetrics how easily a plausible technology can become indispensable before its promised clinical benefit is established.
AI gives us a second chance.
This time, we should ask for the outcomes first.
One important substantive change: I would not retain the sentence that cerebral palsy “has sat near 2 per 1,000 births the whole time.” The broader point is defensible, but that formulation implies a longitudinal constancy that is harder to support cleanly across changing diagnostic definitions and populations. Likewise, I replaced the original ultrasound statement with the stronger 2026 peer-reviewed evidence now available. (PubMed)


