Series
Building and stress-testing clinical AI in public.
A build-in-public series on medical diagnosis models. The name is a double entendre: the differential diagnosis — the ranked list of candidate diseases — and the derivative — how much the output should move when the input changes. The question driving the series: when the evidence changes, does the model's conclusion change in a clinically coherent way? Each post documents a real experiment — perturbing cases, ablating findings, injecting counterevidence — and what it reveals about whether these models reason from evidence or recite patterns.
Now
Moving off clean vignettes. Testing whether a specialized system can beat a frontier model on the messy parts of a diagnostic odyssey: long fragmented records, contradictory evidence, source provenance, multimorbidity, analogous-case retrieval, and the value of the next missing test.
A frontier model ranked the correct diagnosis first on a masked neuromuscular case. Then I started removing evidence — and its confidence didn't move the way it should. Part 1 of The Differential.
Jul 29, 2026 · 4 min read
I repeated every perturbation condition three times. The leading diagnosis was highly reproducible; the ordering and scoring below it was not — which killed my original product hypothesis. Part 2 of The Differential.
Jul 29, 2026 · 9 min read