Three maternal-fetal medicine specialists read the same fetal heart rate tracings. On the category we treat as an emergency, their agreement was kappa 0.0. That is the test we have used as the backbone of antenatal surveillance for fifty years.
A woman at 39 weeks reports reduced fetal movement. She gets a non-stress test. Twenty minutes of tracing, a few accelerations, no decelerations. One clinician calls it reactive and sends her home. Another looks at the same strip and wants another twenty minutes. Neither is wrong, because there is no measurement being made. There is only a reading.
The non-stress test asks whether the fetal heart rate accelerates.
Reactive means two accelerations in twenty minutes.
That definition sounds objective.
The act of applying it is not.
When Blackwell and colleagues gave 154 fetal heart rate segments to three maternal-fetal medicine specialists and asked them to assign NICHD categories, interobserver reliability was moderate at kappa 0.45. For category III, the tracings we treat as the emergency, agreement was kappa 0.0. Not poor. Zero. The disagreement came down to whether variability was absent or minimal. These were subspecialists, not residents.
The finding is not an outlier. When six clinicians applied the 2015 FIGO guidelines to 151 tracings, proportions of agreement for overall classification ran from 0.54 to 0.67, and experience made no difference. That last detail deserves more attention than it gets. If more years on labor and delivery do not improve agreement, the variability is not a training problem. It is built into the task.
Most of this literature is intrapartum, because that is where the lawsuits are.
The antepartum reliability literature is thinner, which is itself part of the indictment. We have never seriously measured the reproducibility of a test we order millions of times a year.
An objective alternative has existed since the 1980s. Geoffrey Dawes and Christopher Redman built a computerized analysis at Oxford that returns a binary answer, criteria of normality met or not met, derived from a database of well over 100,000 traces and their outcomes. Its central output is short-term variation, the millisecond-level beat-to-beat difference that no eye can compute. TRUFFLE used a short-term variation of 3.5 milliseconds as an intervention trigger in early-onset growth restriction. Bhide and colleagues then examined 14,025 computerized assessments and found criteria unmet in roughly one of every sixteen.
Stillbirth was almost nine times more frequent in that group, odds ratio 8.78, 95% confidence interval 4.28 to 18.02. Exclude the cases with low short-term variation and the signal holds, odds ratio 7.62.
So here is the state of the evidence. We have a subjective test with documented agreement problems in the hands of subspecialists, and an objective test with a strong prognostic signal across 14,025 pregnancies.
The obvious next step is a trial. Now look at what the profession actually did.
The Cochrane review of antenatal cardiotocography includes six trials and 2,105 women in total. The comparison that matters, computerized versus visual interpretation, rests on two studies and 469 women. It showed a reduction in perinatal mortality, 0.9 percent versus 4.2 percent, risk ratio 0.20 with a confidence interval of 0.04 to 0.88 that nearly touches one. A 2021 systematic review found three randomized trials, 497 women, and a single antenatal stillbirth across all of them.
Four hundred sixty-nine women. That is the entire randomized evidence base for whether objective interpretation of the fetal heart rate saves babies.
For comparison, the ARRIVE trial randomized 6,106 low-risk nulliparous women to answer a question about the timing of induction, and the field changed its practice within two years. We found the money, the sites, and the will for that. We have never found them for this.
The consequence lands on the patient. A woman sent home after a reactive non-stress test believes a measurement was made.
What actually happened is that one clinician looked at a strip and formed an impression another clinician might not have shared. She was never told that. It is not in any consent conversation I have ever heard.
Conclusion
The algorithm is not the problem. Dawes-Redman received FDA clearance in March 2025, so in the United States the regulatory excuse is gone, and what remains is a purchasing decision and a research agenda nobody has demanded. The honest reading is this.
We built the backbone of antenatal surveillance out of a subjective judgment.
We have known for at least fifteen years that experienced subspecialists disagree about it, including on the category we call an emergency. And we never ran the trial that would tell us whether the objective version does better.
That is not a failure of technology or of regulators.
It is our failure.
A profession that can randomize six thousand women to settle a question about induction timing can randomize enough women to settle this one.
If you order non-stress tests, ask what your agreement rate with your partner would be on the last ten you read.
Then ask why nobody has ever measured it. ObGyn Intelligence is free, and it stays independent because paid subscribers keep it that way.
References
1. Blackwell SC, Grobman WA, Antoniewicz L, Hutchinson M, Gyamfi Bannerman C. Interobserver and intraobserver reliability of the NICHD 3-Tier Fetal Heart Rate Interpretation System. Am J Obstet Gynecol. 2011;205(4):378.e1-5. doi:10.1016/j.ajog.2011.06.086
2. Rei M, Tavares S, Pinto P, Machado AP, Monteiro S, Costa A, et al. Interobserver agreement in CTG interpretation using the 2015 FIGO guidelines for intrapartum fetal monitoring. Eur J Obstet Gynecol Reprod Biol. 2016;205:27-31. doi:10.1016/j.ejogrb.2016.08.017
3. Bhide A, Meroni A, Frick A, Thilaganathan B. The significance of meeting Dawes-Redman criteria in computerised antenatal fetal heart rate assessment. BJOG. 2024;131(2):207-212. doi:10.1111/1471-0528.17464
4. Grivell RM, Alfirevic Z, Gyte GML, Devane D. Antenatal cardiotocography for fetal assessment. Cochrane Database Syst Rev. 2015;(9):CD007863. doi:10.1002/14651858.CD007863.pub4 [issue number pending RefVerify]
5. Baker H, Pilarski N, Hodgetts-Morton VA, Morris RK. Comparison of visual and computerised antenatal cardiotocography in the prevention of perinatal morbidity and mortality. A systematic review and meta-analysis. Eur J Obstet Gynecol Reprod Biol. 2021. doi:10.1016/j.ejogrb.2021.05.048 [volume and pages pending RefVerify]
6. Lees CC, Marlow N, van Wassenaer-Leemhuis A, Arabin B, Bilardo CM, Brezinka C, et al. 2 year neurodevelopmental and intermediate perinatal outcomes in infants with very preterm fetal growth restriction (TRUFFLE): a randomised trial. Lancet. 2015;385(9983):2162-72. [pending RefVerify]
7. Grobman WA, Rice MM, Reddy UM, Tita ATN, Silver RM, Mallett G, et al. Labor induction versus expectant management in low-risk nulliparous women. N Engl J Med. 2018;379(6):513-523. doi:10.1056/NEJMoa1800566 [pending RefVerify]
8. Huntleigh Healthcare. FDA 510(k) clearance granted for Dawes-Redman CTG Analysis. Press release, 24 March 2025.


