Something extraordinary has happened. Medicine should be paying attention.
I have spent much of my professional life in obstetrics, maternal-fetal medicine, clinical research, and medical ethics. I have watched medicine make remarkable progress. I have also watched us repeatedly fail to answer questions that are fundamental to the health and survival of women and their children.
We have more publications, larger databases, more sophisticated statistical methods, and more clinical guidelines than at any previous point in medical history.
Yet some of our most important questions remain unanswered.
Why does preeclampsia develop in some pregnancies but not others? Why does spontaneous preterm labor begin? Why do some fetuses die unexpectedly? Why can we still not reliably prevent cerebral palsy? Why does endometriosis cause devastating pain in some women while others have extensive disease and few symptoms?
For decades, we have treated these questions as extraordinarily difficult problems that might eventually yield to conventional medical research.
I now believe we need to consider a more ambitious possibility: What if the problem is not simply that medicine lacks information, but that we have not yet developed the intellectual tools needed to understand the information we already possess?
On October 6, 2026, OpenAI released a collection of mathematical research findings produced by an advanced internal AI model.
The initial release contained 722 manuscripts covering 372 families of related results. These included investigations of difficult mathematical questions, some of which had resisted conventional approaches. They were placed in a public GitHub repository with supporting proofs, research materials, and mechanisms for revisions.
Not every proposed mathematical result has been independently verified. Some have formal machine-checkable proofs, while others require critical examination. Their significance is still being assessed.
Nevertheless, the implications are potentially enormous.
We are no longer discussing AI exclusively as a tool that summarizes published research, writes manuscripts, or answers questions.
We are witnessing AI being used to generate new scientific arguments and proposed solutions to problems that previously lacked accepted answers.
That distinction may eventually become one of the most important developments in the history of medicine.
The difference between answering questions and discovering answers
Most physicians currently use artificial intelligence as a sophisticated information system.
We ask what the evidence says about a treatment. We request summaries of randomized trials. We ask for differential diagnoses, analyze datasets, or generate drafts of research papers.
These are valuable applications. But they primarily involve reorganizing, interpreting, or applying existing knowledge.
Scientific discovery requires something more.
It requires recognizing relationships that have not previously been identified, developing hypotheses that humans have not considered, finding weaknesses in accepted assumptions, and designing ways to distinguish competing explanations.
That is why the mathematical developments are so important.
They suggest that advanced AI systems may be approaching a level of research capability at which they can contribute genuinely new intellectual work.
Mathematics is particularly suited to this development because a mathematical proof can, in principle, be evaluated against explicit logical rules.
Medicine is different.
A mathematical theorem can be true within a defined formal system. A proposed explanation for preeclampsia cannot be established by logic alone. It must survive biological investigation, clinical validation, and testing in actual patients.
But this is not an argument against the application of AI to medicine.
It is an argument for combining AI’s potential discovery capabilities with rigorous clinical research.
Imagine an AI system capable of integrating molecular biology, placental pathology, maternal physiology, genetic information, clinical trajectories, and decades of epidemiological research.
Now imagine that system generating entirely new explanations, identifying overlooked interactions, and proposing experiments that could distinguish between competing mechanisms.
The great opportunity is not to replace medical research with artificial intelligence. It is to expand the number, quality, and originality of the questions medical research can investigate.
The ten greatest mysteries in obstetrics and gynecology that AI should help us solve
This is my proposed list. It is not an established ranking of scientific importance. These ten questions reflect the burden of disease, the limits of current understanding, and the potential for discovery to change clinical care.
1. Why does preeclampsia develop, and can we prevent it before the disease begins?
Preeclampsia remains one of the most consequential complications of pregnancy.
We understand important elements of placental development, angiogenic imbalance, endothelial dysfunction, and maternal susceptibility. We can identify some patients at increased risk, and aspirin reduces the risk in selected populations.
But we still cannot fully explain why one patient develops severe early-onset preeclampsia while another patient with seemingly similar characteristics remains unaffected.
Could AI identify previously unrecognized biological pathways or distinct disease mechanisms hidden within the broad clinical diagnosis of preeclampsia?
The ultimate objective should be prevention based on an understanding of disease mechanisms, not merely earlier recognition of hypertension and organ dysfunction.
2. What actually initiates spontaneous preterm labor?
After decades of research, spontaneous preterm birth remains a major unresolved problem.
Infection, inflammation, cervical remodeling, decidual activation, uterine stretch, and other pathways are implicated. Yet no single mechanism explains all cases.
We often approach preterm birth as though it were one disease. It may instead be a final common clinical event arising from multiple biologically distinct processes.
AI could help identify those processes, determine how they interact, and discover which pregnancies are likely to respond to specific preventive interventions.
The goal is not simply to predict who might deliver early.
The goal is to understand why labor begins prematurely and to develop effective ways to prevent it.
3. Why do apparently healthy fetuses die unexpectedly before birth?
Stillbirth remains one of obstetrics’ most devastating outcomes.
Some deaths are associated with placental disease, fetal anomalies, infection, or maternal illness. Others remain unexplained even after careful investigation.
Our inability to identify every fetus at risk is a fundamental limitation of prenatal care.
Could AI integrate longitudinal fetal growth, placental biology, maternal characteristics, Doppler findings, and other physiological signals to recognize previously invisible patterns of deterioration?
Could it help distinguish fetuses who would benefit from earlier delivery from those who would be harmed by unnecessary intervention?
The critical challenge is not producing another risk score.
It is discovering mechanisms that permit preventable fetal deaths to be prevented without creating equally serious harms through unnecessary early delivery.
4. Why does fetal or neonatal brain injury occur, and how can we reliably prevent it?
Hypoxic-ischemic injury, neonatal encephalopathy, and cerebral palsy are often discussed together, although they are not interchangeable conditions.
Some injuries originate before labor. Others arise during labor, delivery, or the neonatal period. Their mechanisms are heterogeneous, and the timing of injury is frequently uncertain.
We still lack sufficiently accurate methods to reconstruct the development of many brain injuries and to identify effective preventive opportunities.
AI might integrate fetal monitoring, maternal disease, placental pathology, neonatal findings, neuroimaging, and long-term developmental outcomes.
Could it identify causal pathways that current clinical classifications obscure?
The central objective must be preventing neurological injury, not merely predicting or retrospectively assigning responsibility for an adverse outcome.
5. Why do some pregnancies develop fetal growth restriction while others do not?
Fetal growth restriction is not simply a problem of identifying a fetus below a statistical weight threshold.
Some constitutionally small fetuses are healthy. Others have impaired growth because of placental dysfunction, maternal disease, fetal conditions, or multiple interacting causes.
Current approaches remain imperfect in separating normal biological variation from pathological growth restriction.
AI could potentially integrate individual growth trajectories, placental biomarkers, fetal hemodynamics, and maternal physiology to identify more meaningful biological phenotypes.
The goal should be to distinguish fetuses who need intervention from those who do not.
Can we finally move beyond population-based fetal size thresholds toward a reliable understanding of individual fetal growth and fetal well-being?
6. Why does endometriosis develop, and why does disease severity correlate so poorly with symptoms?
Endometriosis affects millions of women, yet major questions about its origins, progression, pain mechanisms, and treatment response remain unresolved.
The extent of visible disease does not reliably explain the severity of pain or infertility.
Some patients experience profound disability. Others with substantial anatomical disease have relatively limited symptoms.
AI-assisted research could help integrate genetics, immune function, hormonal signaling, lesion characteristics, neural pathways, and clinical phenotypes.
Perhaps endometriosis, like many obstetric syndromes, includes biologically distinct disorders that we have grouped under a single diagnosis.
The challenge is to explain the disease sufficiently well that treatment can target its mechanisms rather than repeatedly managing its consequences.
7. Why do some pregnancies end in miscarriage despite apparently normal embryos?
Chromosomal abnormalities explain an important proportion of early pregnancy losses.
But not every miscarriage can be attributed to aneuploidy, and repeated losses can occur despite extensive evaluation.
Embryonic development, implantation, uterine biology, immune regulation, maternal health, and placentation are extraordinarily complex.
AI may help identify combinations of biological abnormalities that escape conventional diagnostic categories.
Can we distinguish losses that are biologically unavoidable from those that could be prevented?
A major advance would be moving from broad, often inconclusive investigations toward validated, mechanism-specific prevention.
8. Why does implantation fail, and what determines whether a healthy embryo becomes an ongoing pregnancy?
Assisted reproductive technology has transformed infertility care, but implantation remains incompletely understood.
Even when embryo quality appears favorable, successful implantation and ongoing pregnancy cannot be guaranteed.
The interaction between the embryo and endometrium involves coordinated molecular, cellular, hormonal, and immune processes.
AI could help generate hypotheses about which interactions determine successful implantation and early placental development.
It might also identify which commonly used tests and interventions have little genuine clinical value.
The purpose should not be to invent more expensive fertility technologies. It should be to discover which biological mechanisms actually determine reproductive success.
9. Why do racial disparities in maternal and perinatal outcomes persist, and which interventions will reduce them?
In the United States, major disparities persist in maternal mortality, severe maternal morbidity, preterm birth, and cesarean delivery.
The reasons are complex and cannot responsibly be reduced to a single explanation.
Patients differ in exposure to social and environmental risks, underlying health conditions, access to care, hospital resources, referral patterns, and the care they receive.
For cesarean delivery, for example, differences between hospitals, within hospitals, and across clinically comparable patient groups must be examined separately.
AI could help analyze these complex pathways, identify where disparities arise, and generate testable hypotheses about effective interventions.
But prediction is not causation. An algorithm that identifies a disparity has not explained its cause, and an algorithm trained on inequitable care may reproduce existing inequities.
Our goal should be causal understanding and measurable improvements in outcomes, not simply increasingly sophisticated descriptions of unequal results.
10. Why do we perform so many cesarean deliveries, and how can we identify which ones are truly necessary?
Cesarean delivery is one of the most consequential interventions in obstetrics. It can be lifesaving, but it also creates immediate and future maternal risks.
Yet substantial variation in cesarean rates remains across hospitals, populations, and clinical practices.
The Robson Ten-Group Classification System offers an important framework for examining this variation using defined obstetric populations.
But classification is only the beginning.
We need to determine which clinical decisions produce better outcomes, which practices increase intervention without benefit, and how patient characteristics, hospital organization, and provider behavior influence delivery decisions.
Could AI combine Robson group data with longitudinal maternal and neonatal outcomes, clinical circumstances, and hospital-level characteristics to identify more effective approaches?
That would require careful causal studies, not simply an algorithm trained to reduce cesarean rates.
The objective is not fewer cesareans at any cost. It is the right delivery for the right patient, supported by evidence of better outcomes.
Why medicine may be an even more consequential frontier than mathematics
A mathematical discovery can profoundly change our understanding of the world.
A medical discovery can change whether a mother survives childbirth, whether a fetus is born alive, whether a newborn develops neurological injury, or whether a woman experiences years of preventable pain and infertility.
That is what makes the application of advanced AI to medicine so compelling.
However, medicine faces barriers that mathematics does not.
Our data are fragmented. Diagnoses are often imprecise. Clinical outcomes reflect biological mechanisms, treatment decisions, social conditions, and health system organization.
Much of our published research is observational. Associations are repeatedly mistaken for causes. Important negative studies are less visible than positive findings. Research populations are frequently selected, and results may not apply equally across settings.
AI can magnify these problems if deployed carelessly.
But advanced AI could also help expose them.
It could identify inconsistencies across thousands of studies, detect hidden assumptions, challenge accepted explanations, generate alternative causal models, and propose research designs that are more informative than another retrospective analysis of the same question.
The most valuable application of AI may ultimately be its ability to tell us that the questions we have been asking are incomplete or incorrectly framed.
Can we solve these ten mysteries soon?
Perhaps some of them. But we must distinguish accelerated discovery from proven clinical benefit.
I do not believe the mathematical release justifies predicting that all ten mysteries will be solved within five years, ten years, or any other specified period.
Medicine is not a formal mathematical system. Human biological discoveries require experiments, replication, external validation, and prospective studies. Effective treatments must demonstrate benefits that outweigh their harms.
Even a brilliant AI-generated hypothesis remains a hypothesis.
But I believe something important has changed.
Until recently, the limiting factor in many fields was our ability to generate and evaluate new scientific explanations. Modern AI may substantially expand that capacity.
We should begin developing research programs in which AI systems investigate major unanswered clinical questions, while independent physicians, biologists, statisticians, and ethicists critically assess the results.
Every proposed discovery should be transparent enough to evaluate. Every analysis should identify its data sources and assumptions. Every important causal claim should face an appropriate empirical test.
We should not confuse an impressive model output with scientific truth.
But we should also not confuse the necessary demand for verification with an argument against pursuing discoveries that could transform medicine.
Medicine needs its own version of OpenAI’s mathematical challenge
The most important lesson from the mathematics release may not be any individual proof.
It may be the research model itself.
OpenAI assembled a large collection of difficult open problems, applied an advanced AI system, generated proposed solutions, and placed the results in a publicly inspectable repository.
Why should medicine not develop a comparable initiative?
Imagine a publicly accessible collection of the hundred most important unanswered questions in obstetrics and gynecology.
For each question, investigators could document what is known, what remains uncertain, which hypotheses compete, which datasets are available, and what evidence would be required to establish a meaningful answer.
AI systems could generate candidate explanations and research approaches. Independent clinical experts could evaluate them. Research groups could test the strongest proposals. Negative results, corrections, and failed hypotheses could remain publicly accessible rather than disappearing into unpublished files.
Traditional peer-reviewed journals would still have an important role, but they would no longer need to be the only mechanism through which scientific findings become visible.
Such a system could make research more transparent, cumulative, collaborative, and open to correction.
I have begun thinking about precisely this approach through an Obstetric 100 project, a proposed collection of one hundred major problems that deserve systematic investigation.
The aim is not one hundred more review articles.
It is one hundred clearly defined scientific challenges, each connected to evidence, competing explanations, testable hypotheses, and potential clinical outcomes.
AI could become an important participant in this process, but never an authority beyond independent scientific scrutiny.
The ethical obligation to pursue discovery
As an obstetrician and medical ethicist, I believe this development raises a question of professional responsibility.
We rightly insist that physicians should not use AI systems that have not been appropriately validated for a clinical purpose.
That principle remains essential.
But there is another responsibility.
If emerging AI research methods can help us identify preventable causes of fetal death, maternal complications, infertility, and neurological injury, do we not also have an obligation to investigate those capabilities?
We cannot establish that obligation by assuming AI will succeed. We can, however, recognize the ethical importance of rigorously evaluating methods that might substantially improve patient outcomes.
The choice is not between unrestricted AI use and rejecting AI altogether.
The responsible choice is ambitious research combined with demanding standards of evidence.
I do not want an AI system to declare that it has solved preeclampsia.
I want it to identify an explanation nobody has considered, make that explanation scientifically testable, and help researchers discover whether it is true.
I do not want an algorithm to decide when every patient should deliver.
I want research that allows us to understand more precisely which patients benefit from delivery and which benefit from continuing pregnancy.
I do not want AI to replace physicians or scientists.
I want AI to help physicians and scientists discover what we have been unable to discover on our own.
The next great scientific breakthrough may not come from a laboratory alone
The OpenAI mathematics release does not prove that artificial intelligence can solve medicine’s greatest mysteries.
But it gives us a compelling reason to reconsider the boundaries of scientific discovery.
For too long, much of medical research has advanced through incremental extensions of established methods. Those methods remain indispensable, but they need not define the limits of what comes next.
We now have an opportunity to unite clinical experience, biological experimentation, high-quality data, causal inference, and advanced artificial intelligence in a more powerful discovery process.
Some AI-generated hypotheses will fail. Some purported discoveries will prove incorrect. Others may simply rediscover what we already know.
But a few may open scientific pathways that decades of conventional investigation have missed.
That possibility is too important to dismiss.
My hope is that October 2026 will be remembered not merely as a remarkable moment in the history of artificial intelligence and mathematics, but as a turning point in how medicine approached its most difficult unanswered questions.
The greatest achievement of medical AI will not be writing better clinical notes, producing faster literature reviews, or generating more research manuscripts.
It will be helping us prevent deaths and diseases that we currently cannot prevent.
And if that future is possible, then the question for obstetrics and gynecology is no longer simply what AI can do today.
It is whether we are prepared to ask it the questions that matter most.
Amos Grünebaum, MD
MD Intelligence | Evidence • Insights • Patient Safety
What would you add to the ten greatest unanswered questions in obstetrics and gynecology? And which should an AI-assisted research initiative tackle first?


