ENFR
Insights

Where the answers come apart

Two systems agree on the case that was randomised and diverge on the one standing in front of you.

An NIHSS of 4 with a left M1 occlusion, thrombolysis already running: the drug will open about a fifth of proximal M1 occlusions, closer to half in patients this mild, and the question of whether to move to the angiography suite is still open when the infusion finishes. I have put cases like this one to a general-purpose model and to a corpus of curated, source-attributed evidence. On the ordinary anterior-circulation occlusion inside the randomised window they return the same thing, which is the least useful output either of them produces; the points worth having start to arise as the case gets harder.

On the Angels Stroke Heroes podcast in August I put the general position this way: "If you say I'm going to use AI to solve all the problems, you're out of your mind. If you say I'm not using AI, well, you've been left behind in 2020." I still hold it. With the honeymoon over, the useful work is describing precisely where these systems stop agreeing with each other, because that is where a physician has to decide which one he believes.

They stop agreeing in two different ways, and the first is the patient no trial enrolled. An 81-year-old woman with a prestroke mRS of 3, a distal basilar occlusion and a last known well of nine hours falls outside every randomised envelope that exists: ATTENTION required a premorbid mRS of 0 at her age, and BAOCHE did not enrol above 80. I do not know what the right answer is for her. I do know which trials excluded her and on which criterion, and that is a different kind of not-knowing from the one a fluent paragraph of synthesis produces.

The same thing happens at the other end of severity. A large core with an unknown time of onset now sits inside an evidence base that six randomised trials rebuilt in three years, and a trial-level Bayesian meta-regression across them puts the absolute benefit below 0.10 at about ten hours from onset and below 0.03 at about eighteen, the second an extrapolation its authors flag. When nobody knows when the onset was, that curve has no argument to offer, and the honest output says so rather than choosing a time branch for me.

The second way is the case where the published evidence disagrees with itself. MILD-MT randomised in favour of immediate thrombectomy in mild deficits, a result presented in 2026 and not yet published in full; a meta-analysis of eleven observational cohorts I published with colleagues found no functional advantage and roughly three times the symptomatic haemorrhage, with all the selection that observational treatment data carry. A system that returns one of those and not the other has made an editorial decision on my behalf without telling me it made one.

A curated corpus is an editorial act too. I decide what is admitted and what is refused, and I should be held to that; a general model given retrieval can also cite a trial, and the question is then which trial it retrieved and whether it retrieves the same one tomorrow. The line I can defend runs between an answer whose sources, admissions and refusals are on the page where a colleague can dispute them and one whose are not, and between an answer that is the same tomorrow and one that is not.

Reproducibility is not validity; a corpus reproducibly returns whatever its curator admitted. What it buys is an audit. An answer that changes between two runs of the same question cannot be checked, and an acute stroke decision has to be reconstructable months later in front of colleagues who were not in the room. Stable comes before better, because "better" is not measurable until the same input reliably returns the same answer.

Attribution works from the other direction. An answer that shows its sources can be argued with — you open the trial and find the enrolment criterion that excludes your patient. An answer with nothing behind it can only be trusted or ignored, and neither of those is a clinical act.

The Musée des Arts et Métiers keeps a room of nineteenth-century machines that very nearly flew, and what I took from it is that knowledge is largely the inhibition of wrong connections. At the margin of a stroke decision that is most of the job: the plausible link between a trial and the patient on the table is the one that has to be tested and, more often than not, refused. Fluency is extremely good at making that link and has no mechanism for refusing it.

None of this decides anything, and it is not built to. What I am assembling is navigational guidance, not a medical device: it organises the evidence for a decision and leaves the decision where it already sits. The physician who signs the note carries the consequence, and the only version of this technology worth having is the one that makes that signature easier to defend.

Angels Stroke Heroes, episode 17, "Dr Apostolos Safouris | Lessons From History", produced by The Angels Initiative, published 20 August 2026 — https://podcasters.spotify.com/pod/show/jayd38/episodes/Dr-Apostolos-Safouris--Lessons-From-History-e3nl8sh

The three worked cases — LiveTextbook