Rob Parisien, MD, FAAOS
Theme Nine

When the instrument is a machine

What kind of witness is a machine, and on which population was it right?

The organising claim of this theme is that artificial intelligence raises almost no epistemological problem the preceding entries have not already raised. Shortcut learning is a reference class failure. Automation bias is a calibration failure. The argument about explainability is the argument about mechanism and outcome. What is new is that the errors are distributed differently from ours and arrive in a confident register, which makes them harder for us to catch than our own. The theme opens with the best evidence any of this has, because a reading list assembled out of failures would flatter our scepticism and teach nothing.

40

Mammography Screening with Artificial Intelligence (MASAI)

Lancet Oncol 2023, Lancet Digit Health 2025, Lancet 2026 · Lång K, Hernström V, Gommers J et al. · Randomised, controlled, population-based screening-accuracy trial, n=105,934

What it does: Randomises women in the Swedish national screening programme to AI-supported screen reading or standard double reading. The AI triaged examinations to single or double reading and marked suspicious findings; radiologists still read every case. Cancer detection was 6.4 against 5.0 per 1000 (ratio 1.29) with no increase in false positives. The primary endpoint, reported in 2026, was interval cancer rate: non-inferior at 0.88, with sensitivity 80.5% against 73.8% (p=0.031) and specificity identical at 98.5%. Screen-reading workload fell by 44%.

Epistemological angle: This is the entry that holds AI to the standard the rest of this collection demands of everyone else, and it is worth being explicit that it meets that standard where most of the sham-tested procedures in theme one did not. Randomised, population-based, pre-registered, primary endpoint pre-specified and reported on a hard outcome rather than a surrogate. Note what was actually tested: not the machine replacing the radiologist but the two together against the standard of care. The comparator matters as much as the result — Swedish double reading is a demanding benchmark, and the intervention beat it on sensitivity while halving the human reading burden. What licenses the result is also worth naming: the model's training distribution and its deployment population were the same, which is precisely the condition entry 42 shows to be absent when these systems fail.
My take: I include this first in the theme deliberately. It would be easy to assemble an AI reading list out of failures, and the result would be a collection that flatters our scepticism and teaches nothing. The honest position is that the best available evidence for AI-supported reading is stronger than the evidence for a good deal of what I do in theatre, and I would rather say so plainly than bury it.
Lancet Oncol 2023;24(8):936–944 · PMID 37541274 · doi:10.1016/S1470-2045(23)00298-X · Lancet Digit Health 2025;7(3):e175–e183 · PMID 39904652 · doi:10.1016/S2589-7500(24)00267-X · Lancet 2026;407(10527):505–514 · PMID 41620232 · doi:10.1016/S0140-6736(25)02464-X
41

An Algorithmic Approach to Reducing Unexplained Pain Disparities in Underserved Populations

Nat Med · 2021 · Pierson E, Cutler DM, Leskovec J, Mullainathan S, Obermeyer Z · Deep learning on knee radiographs

What it does: Trains a model on knee radiographs to predict the patient's reported pain rather than to reproduce a radiologist's severity grade. Standard radiographic grading accounted for 9% of unexplained disparity in pain between served and underserved patients; the algorithmic prediction accounted for 43%. The authors conclude that much of the unexplained pain arises from features within the knee that standard radiographic measures do not capture, and note that since severity measures drive treatment decisions, this bears directly on access to arthroplasty.

Epistemological angle: An orthopaedic classification in use since 1957 was shown to be lossy by a system with no stake in it. That is worth separating from the usual framing. The finding is not that the algorithm is cleverer; it is that the radiographic grade discards information the knee contains, and that what it discards is not randomly distributed. Read against entry 47, this is the same lesson from the opposite direction: an ordinal category imposed on a continuous structure loses signal, and the loss is invisible from inside the category because the category is what you are looking with. It is also the collection's clearest case of a machine functioning as a genuinely third kind of cognizer — not faster at our task, but sensitive to something our task was never set up to see.
My take: This is the paper that changed my mind about what these systems are for. I had thought of them as accelerating what we already do. Here the value came from asking a different question — predict the symptom, not the grade — and the answer implied that our grading scheme was the problem. I would like to see this design repeated in the shoulder and the spine, and I do not know why it has not been.
Nat Med 2021;27(1):136–140 · PMID 33442014 · doi:10.1038/s41591-020-01192-7
42

Deep Learning Predicts Hip Fracture Using Confounding Patient and Healthcare Variables

npj Digit Med · 2019 · Badgeley MA, Zech JR, Oakden-Rayner L et al. · 17,587 pelvic radiographs

What it does: Trains models to classify hip fracture and, separately, five patient traits and fourteen hospital process variables. All twenty could be predicted from the radiograph, with the best performance on scanner model (AUC 1.00), scanner brand (0.98) and whether the order was marked priority (0.79). Fracture prediction reached AUC 0.78 from the image alone and 0.91 with patient and process data added. On a test set balanced across patient and process variables the model performed at chance, AUC 0.52.

Epistemological angle: The reference class problem of entry 17 in its purest mechanical form, and the reason that entry belongs upstream of this whole theme. A trained model's reference class is its training distribution, and nothing in the model selects that class — the data do, silently, through whatever happens to covary with the label. Here the class turned out to be something like "pelvic radiographs from this scanner, at this hospital, ordered with this priority flag." The 0.52 is the number to carry: strip the confounders and the fracture detector detects nothing. Hájek's point that the choice cannot be eliminated but only made visible applies exactly, and it is why the deployment question for these systems is not how accurate they are but on which population that accuracy was estimated, and whether the patient in front of you is in it.
My take: Set beside entry 40, which works, this stops being a cautionary tale and becomes the criterion. MASAI's population and its training population were the same. This model's were not, and nobody knew because nobody had asked what class it had learned. That is a question a surgeon can ask of a vendor without any machine learning at all, and it is the only one I would insist on.
npj Digit Med 2019;2:31 · PMID 31304378 · doi:10.1038/s41746-019-0105-1
43

Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance

Radiology · 2023 · Dratsch T, Chen X, Rezazade Mehrizi M et al. · Prospective experiment, 27 radiologists, 50 mammograms

What it does: Radiologists of varying experience read mammograms with the aid of a purported AI system that supplied a deliberately incorrect BI-RADS category for twelve of forty cases. Correct assessments fell from 79.7% to 19.8% among inexperienced readers, from 81.3% to 24.8% among the moderately experienced, and from 82.3% to 45.5% among the very experienced.

Epistemological angle: The counterpart to entry 40, and the reason a favourable trial does not settle the question. MASAI shows what a well-matched system adds to a reader; this shows what a wrong one subtracts, and the subtraction is larger than the addition. The epistemological content is about testimony. A machine's output enters the reader's judgment not as one datum among several but as an anchor, and the very experienced group — down thirty-seven points — demonstrates that expertise is protection rather than immunity. Read with entry 19: the failure is in the same family as the anchoring effects that moved surgeons' answers by 22 points, except that here the anchor arrives with institutional authority attached.
My take: The design is artificial and the effect size is therefore an upper bound; nobody deploys a system that is wrong 30% of the time. What survives the artificiality is the direction and the gradient across experience, and those are enough. If I am going to use these tools, I should want to know my own susceptibility, and I have no way to measure it.
Radiology 2023;307(4):e222176 · PMID 37129490 · doi:10.1148/radiol.222176
44

Endoscopist Deskilling Risk After Exposure to Artificial Intelligence in Colonoscopy

Lancet Gastroenterol Hepatol · 2025 · Budzyń K, Romańczyk M, Kitala D et al. · Multicentre observational study, 1,443 patients, four Polish centres

What it does: Compares the adenoma detection rate of standard, unassisted colonoscopy in the three months before AI detection tools were introduced against the three months after. Detection on unassisted procedures fell from 28.4% to 22.4%, an absolute difference of 6.0% (adjusted odds ratio 0.69). The comparison is of the endoscopists' own unaided performance, before and after they became habituated to assistance.

Epistemological angle: This is the theme's most important entry for anyone thinking about how to integrate these systems, and it is not an argument against them. Entry 40 establishes that the combination outperforms the human alone. This establishes that the human inside the combination may not be the human who entered it. Both can be true, and taken together they reframe the question from whether to adopt to what to measure afterwards — the unaided performance of the assisted clinician is a quantity nobody was tracking, and it turns out to move. The design cannot establish causation and the authors do not claim it does; it is a before-and-after comparison with all the vulnerabilities that implies.
My take: Six percentage points of adenoma detection is not a small number, and it was found by people who were running an AI trial and thought to look sideways. That is the part I find instructive. The effect was available to be measured for as long as these tools have been deployed, and it was measured once, late, and by accident of study design. Whatever the true magnitude, the more durable finding is how nearly we did not look.
Lancet Gastroenterol Hepatol 2025;10(10):896–903 · PMID 40816301 · doi:10.1016/S2468-1253(25)00133-5 · see also the correspondence and authors' reply, Lancet Gastroenterol Hepatol 2025;10(12):1062, PMID 41205617, doi:10.1016/S2468-1253(25)00324-3
45

Ghassemi versus London: Must a Clinical Machine Be Able to Explain Itself?

Lancet Digit Health · 2021 · Ghassemi M, Oakden-Rayner L, Beam AL · and Hastings Cent Rep · 2019 · London AJ · Opposing arguments

What it does: Ghassemi and colleagues argue that current explainability methods do not deliver what is claimed for them — they produce post-hoc rationalisations that can fail silently for the individual patient — and that rigorous internal and external validation is the more direct route to the goals explainability is supposed to serve. London argues that opaque decisions are far more common in medicine than critics acknowledge, and that where causal understanding is thin, the ability to produce and empirically verify results can matter more than the ability to explain how they were produced.

Epistemological angle: Taught as a pair, in the manner of a genuine disagreement between competent people rather than a controversy with a right answer. The dispute is the mechanism-versus-outcome argument this collection has already run twice, transposed. Entry 23 shows a drug that suppressed the mechanism and killed patients; entry 39 shows repairs that fail structurally in patients who do well. In both, the causal story we could tell was disconnected from the result we could verify. London's argument is that this is medicine's normal condition and that demanding explanation from machines while tolerating its absence in ourselves is inconsistent. Ghassemi's is that the inconsistency does not license accepting an explanation we know to be confabulated. Both are correct about something, and the collection cannot resolve it.
My take: London's argument is the uncomfortable one for me, because it is continuous with my own material. I have spent a good deal of this collection arguing that we reason from mechanism when we should reason from outcome, and it would be inconsistent to reverse that the moment the reasoner is a machine. Where I part from him is on what verification requires: entry 42 shows that a model can be verified impeccably on the wrong population, and no amount of outcome data from that population tells you so.
Lancet Digit Health 2021;3(11):e745–e750 · PMID 34711379 · doi:10.1016/S2589-7500(21)00208-9 · and Hastings Cent Rep 2019;49(1):15–21 · PMID 30790315 · doi:10.1002/hast.973
46

Can Artificial Intelligence Pass the American Board of Orthopaedic Surgery Examination?

CORR · 2023 · Lum ZC · Benchmark study, 207 in-training examination questions

What it does: Puts orthopaedic in-training examination questions to a large language model and benchmarks against five years of resident performance data. 47% correct, roughly first-year level, below the tenth percentile of final-year residents — with accuracy falling sharply as questions moved from recall toward application.

Epistemological angle: The right frame is testimony, not competence. The question is not whether the machine is clever but under what conditions its output may be taken on trust, and the answer here is precise: it degrades exactly where clinical judgment lives. A source that is reliable on recall and unreliable on synthesis is a genuinely difficult kind of witness, because its failures arrive in the same confident register as its successes. The specific numbers will date within months. The framing will not.
My take: Much of this collection concerns the ways our calibration, our use of base rates, and our consistency fall short of the ideal. It would be convenient to conclude that machines should therefore take over, and this paper is a caution against that: the machine's errors are differently distributed, not absent, and we are worse at detecting them because they do not look like human mistakes.
Clin Orthop Relat Res 2023;481(8):1623–1630 · PMID 37220190 · doi:10.1097/CORR.0000000000002704

Also worth reading

Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs. PLoS Med 2018;15(11):e1002683 · PMID 30399157 · doi:10.1371/journal.pmed.1002683. The general case behind entry 42: sorting by hospital alone achieved AUC 0.861 because pneumonia prevalence differed between sites by a factor of thirty, and the networks identified the source institution in 99.95% of images.