Writing
The Differential · Part 5
My cross-validation reported 0.61. The pooled out-of-fold number was 0.52. Both were correct, and the gap had nothing to do with the model — it was five folds scoring on five incompatible scales. Part 5 of The Differential.
9 min read
The Differential · Part 4
The text benchmark left me arguing about what counts as a finding. So I moved to pancreatic CT, where a wrong answer is just wrong. The first run scored 0.52 AUROC — and the way it failed was worth more than the score. Part 4 of The Differential.
9 min read
The Differential · Part 3
50 manual extraction runs across two clinical NLP benchmarks. Source grounding was 100% and the unsupported-claim rate was zero — but the evidence units, entity boundaries, and temporal relations fell apart. Part 3 of The Differential.
12 min read
The Differential · Part 2
I repeated every perturbation condition three times. The leading diagnosis was highly reproducible; the ordering and scoring below it was not — which killed my original product hypothesis. Part 2 of The Differential.
9 min read
The Differential · Part 1
A frontier model ranked the correct diagnosis first on a masked neuromuscular case. Then I started removing evidence — and its confidence didn't move the way it should. Part 1 of The Differential.
4 min read
I sat down for a day and attacked my own app. Six categories of finding — and what I'd build into the next project from day one.
4 min read