What happened before this release went public
Before the pages in this release went live, four PMIDs, fourteen DOIs and one trial registration were re-verified at source. The copy came back corrected in three places where it had drifted from the abstracts it cited, and one claim that could not be confirmed was deleted. Twenty-one built pages then sat unpublished until a physician had ruled on every block of text — sixteen recorded rulings, one of them a flat rejection. The page you are reading is the rejected one, rewritten.
That is the method, seen once, on its own output. The rest of this page describes it in general, because a reader who appraises evidence for a living deserves to know how ours is assembled: by a physician working with an AI model — Claude, made by Anthropic — under a written discipline that decides what the model may do, what it may never do, and who rules on every word before it becomes public.
We name the tool plainly, the way we name Leaflet, which draws our maps, or Cloudflare, which serves these pages. Vagueness about how work is done is a poor foundation for a site whose entire argument is that sources should be checkable.
A specification before any build
Nothing substantial is drafted before the decisions are made and written down. A deliverable of consequence starts with a specification: what it is for, whom it serves and whom it does not, what success looks like, and — for every build step — the decision embedded in it, the default that will be taken, and the single word that reverses that default later. The pages you are reading were built from such a document; so was the evidence corpus; so was this paragraph's own rewrite.
The reason is economic. A wrong decision caught in a specification costs one reading; the same decision caught after the build costs the build. Twenty years of stroke-unit protocols teach the same arithmetic — the time to argue about the pathway is before the patient is in it.
One adversarial pass, then a verdict
Every substantial deliverable takes one structured adversarial critique before it ships: an independent second pass whose brief is to argue the strongest case against the work. The critique is then adjudicated line by line — adopted where it names a real defect, held where the original reasoning survives — and the adjudication is recorded with reasons. The specification behind this release came back marked "fix first" and was corrected in three named ways before a single page was built.
The pass runs once. We do not loop a model against its own output, because we have watched iterated self-critique make work worse, and we wrote the observation down. Critique here is a verdict, judged once.
Citations are real or absent
The standing rule, quoted as it is written: "Every PMID, DOI and trial registration is verified against PubMed, ClinicalTrials.gov or the source itself before it appears in anything. Never fabricate; if unverifiable, say so and leave it out." In practice this means every identifier on this site has been fetched at its source, every external link opened live with its title and destination recorded, and every claim that failed the check removed rather than softened into "data suggest".
There is a museum in Paris, the Arts et Métiers, full of nineteenth-century machines that almost worked; we have written before about its real lesson — that knowledge advances as much by learning to inhibit wrong connections as by forming new ones. A reference check is exactly that: an inhibition mechanism. Language models form connections fluently; the discipline is in what gets stopped.
The physician rules on cards
Machine proposals reach the physician as cards: the source quote above, the proposed text below, and four possible rulings — accept, reject, accept with an edit, or hold. The rulings go into decision files with dates. Nothing on this site published itself; every public word passed an explicit ruling, and where a proposal was accepted with an edit, the edit is the physician's hand and the record says so.
The record is a file, not a memory. A decision that lives only in a conversation has a way of un-happening; a decision in a dated file can be audited by anyone who later needs to know why a page says what it says.
Stable before better
The same input must return the same answer before "better" means anything. A system whose output changes between two runs of the same question cannot be evaluated, and an unevaluable system cannot be improved — whatever it is doing, it is not progress. So reproducibility is checked mechanically here: the site's guard scripts are tested by planting deliberate mutations and proving the guards still fail loudly, and the sections of the site that must not change are compared byte for byte after every build. A guard that cannot be made to bite is decoration.
What the model never does here
Claude sees no patient data. The teaching cases on this site were written as fiction from the start, and no identifiable patient reaches the model in any workflow behind this site. Claude publishes nothing on its own; every public word passed a physician's ruling, including these. And Claude makes no clinical recommendation: the systems we build organise evidence around a decision and leave the decision where it already sits — with the physician who signs, and who carries the consequence. The humans choose; the software remembers what they chose.
What this costs, and why we pay it
The discipline is slow and the slowness shows. Pages sit dark for days waiting for a ruling; claims die in verification that would have survived on a faster site; a finished page can be rejected whole and rewritten, as this one was. We pay these costs for a simple reason: in this field a reader's next step after reading may touch a treatment decision, and a page that might be wrong is more expensive than a page that is late.