← Back home

Series

The Differential

Building and stress-testing clinical AI in public.

A build-in-public series on Diagnostic Odyssey — work aimed at patients whose diagnosis has taken years. The name is a double entendre: the differential diagnosis, the ranked list of candidate diseases, and the derivative, how much the output should move when the input changes. Each post is a real experiment designed to kill an assumption before I build on it. So far that has meant perturbing cases to test evidence sensitivity, measuring whether a frontier model can turn clinical narrative into a structure worth reasoning over, and now training on pancreatic CT where a wrong answer can't be argued into a disagreement about definitions. The experiments have retired more of my hypotheses than they have confirmed, which is the point.

Now

The audit came back: the pooled collapse was five folds scoring on disjoint scales, not a weak model. Re-scoring the saved checkpoints across 22 aggregation rules found no case-level signal in the spatial head either. Running the cheapest remaining test — can this pipeline memorise ten cases? If it cannot, the defect is a bug rather than a data-scale problem, and no amount of new architecture would have helped. The 30-case locked test set stays untouched.