OpenAI Forum

Boston Children's on AI that cracked 18 rare-disease cold cases

Katherine Brownstein, Alan Beggs & Suya Shringarpure· Manton Center rare-disease researchers & OpenAI genomics researcher at Boston Children's Hospital
·~39 min·English·OpenAI Forum
LLMReasoningAI CompanyPolicy
TL;DR

A Boston Children's and OpenAI workflow ran o3 deep research over 376 hard rare-disease cases, surfaced evidence that led to 18 new diagnoses, and even proposed a novel gene hypothesis, with human geneticists validating every lead.

01The Human Stakes

The Diagnostic Odyssey

Rare-disease patients wait six to seven years on average for a name, bouncing from specialist to specialist through a search doctors literally call an odyssey.

it unfortunately it's way too common. People often refer to this as a diagnostic odyssey.

Alan Beggs, OpenAI Forum
Key Insight
The panel opened with a patient, Stav, who spent five-plus years chasing an answer as a teenage athlete. The diagnosis was not just closure: knowing his exact gene connected him to a company now developing a drug for his specific condition. The wait is not a delay on the way to care; without the name, the targeted care does not exist.

02The Scale Problem

Needles in a 3-Billion-Base Haystack

Even after filtering, roughly 10,000 candidate variants across up to 1,000 genes survive for a human to weigh, out of three billion DNA bases.

we filter it and remove all the changes that are common in the population, but we're still left with several thousands up to 10,000 different genetic variants.

Alan Beggs, OpenAI Forum
Key Insight
The bottleneck is not that the answer is hidden — it is that no human holds enough of the map. Beggs and Brownstein each know a different slice of the roughly 8,000-9,000 known disease genes; neither knows all of them. An LLM that has read across the whole literature is not smarter than the geneticists, it is simply wider.

03The Time Bottleneck

The Manual Grind

Each promising variant can eat several hours of database and literature checks before an analyst rules it out and starts over on the next one.

there's a huge amount of work that goes into looking at each and every variant. And then sometimes like after several hours, you're like, "Oh, wait, no, no, no. This can't possibly be it." And then you're on to the next one.

Katherine Brownstein, OpenAI Forum
Key Insight
This is why the problem is a good fit for AI and a bad fit for brute force. The scarce resource is not sequencing or storage — it is expert attention, spent linearly, one dead-end variant at a time. Anything that shortens the candidate list pays back in the currency that actually runs out.

04Earning Confidence

Prove It on Solved Cases First

They ran the model on already-solved cases until it scored 80-90%, tuning the prompt against known error modes, before letting it touch an unsolved one.

we could see the sort of the proportion of cases that it got correct increasing to like 80 90%. And then we thought, okay, now this is worth spending a human analyst time on rather than just giving them output that might just be incorrect.

Suya Shringarpure, OpenAI Forum
Key Insight
The discipline here is the transferable part. They did not trust the model because it sounded convincing — they measured it against ground truth they already held, then treated 80-90% as the bar for when its output was worth a human's time. The gate is not accuracy in the abstract; it is accuracy high enough that review is cheaper than starting from scratch.

05Human-in-the-Loop

Nominate, Don't Decide

The model returns two to half a dozen genes, each backed by evidence — a shortlist a geneticist can adjudicate, not a verdict it issues on its own.

It nominates usually anywhere from two to half a dozen potential changes and maybe none of them are the right answer but it very dramatically lowers that universe of information that our human analysts need to think about.

Alan Beggs, OpenAI Forum
Key Insight
The design choice that makes clinicians trust it is that the model attaches its reasoning — this variant, this literature, this phenotype match — instead of returning a bare gene name. That is what turns it from a black box into a research assistant: the human is not asked to believe the model, only to check its evidence. It augments the expert rather than replacing the judgment.

06The Result

376 Cases, 18 New Diagnoses

Across 376 cold cases the workflow surfaced 18 diagnoses — and in one case proposed a genuinely new gene hypothesis no one had connected before.

we get tired looking at papers on PubMed. You know, it's like pages and pages of articles. go to like page four and you think you've done a great job. This was going all the way to page like 30

Katherine Brownstein, OpenAI Forum
Key Insight
The 18 are retrieval wins — links that existed but had not reached this case yet. The quieter breakthrough sits outside that count: the model nominated a gene, S1PR1, and justified it from a 1999 paper deep in the search results. That is synthesis, not lookup — proposing a hypothesis a motivated researcher could have reached with infinite time, in seconds.

07The North Star

The Genome Doesn't Change, the Knowledge Does

Because knowledge grows constantly, old cases can be re-analyzed cheaply every time a new paper lands — turning diagnosis from a one-shot event into a standing process.

the ability to use the models to reanalyze your case essentially as every new paper is published any new piece of knowledge comes in that becomes relevant I think should be quite transformative

Suya Shringarpure, OpenAI Forum
Key Insight
This reframes what a negative result means. Today an unsolved case is closed until someone finds the time to reopen it; here it is simply waiting for the literature to catch up, and an agent watching in the background can reopen it automatically. The unit of progress shifts from human hours to the pace of published science.

08Access and the Human Role

Cheaper Than an MRI

A whole-genome test now runs under $1,000, less than a routine MRI — so the team is building the tool for anyone, not just patients who reach an elite center.

you shouldn't have to be at a tertiary medical center in order to get like this best-in-class like every variant evaluated properly.

Katherine Brownstein, OpenAI Forum
Key Insight
The economics have already flipped — sequencing is a fraction of an MRI insurers pay for without blinking — so the real barrier is access to expert-grade analysis, not the test itself. And the experts are not defensive about handing that analysis to a tool: Beggs is explicit that the work still needs people to interpret and decide. The scarce human judgment is the point, not the casualty.