The Canon — complete
All ten themes, complete, on one page, for reading straight through or printing
Theme One
What surgery looks like without belief
Can a procedure survive being separated from the expectation that surrounds it?
Surgery resisted placebo controls for most of a century, on grounds that were part ethical and part unexamined. The ethical objection was serious and remains unresolved — a sham operation inflicts real harm on a subject with no prospect of benefit to them, and the exchange between Macklin and Freeman in the further reading is worth going to before deciding the field was merely slow. When the controls finally arrived they did something no amount of case-series accumulation could: they separated the operation from the reference class the surgeon had assigned it to. Start here, because the base rate this theme establishes should sit underneath everything that follows.
1
Use of Placebo Controls in the Evaluation of Surgery: Systematic Review
BMJ · 2014 · Wartolowska K, Judge A, Hopewell S et al. · Systematic review, 53 trials
What it does: Collects every placebo-controlled surgical trial then published and asks a simple question of the set. Improvement occurred in the placebo arm in 74% of trials. In just over half, there was no significant difference between the real operation and the sham.
Epistemological angle: This is the base rate for surgery itself, and it belongs in a surgeon's head before any individual outcome study is read. When roughly half of the procedures that have been tested this way fail to beat placebo, an uncontrolled case series reporting good results carries almost no evidential weight — it is exactly what you would expect to see whether or not the operation does anything. The paper also answers the practical objection to sham trials by demonstration: they were done, they were feasible, and the risk was low.
My take: This is a number worth carrying around, and I did not know it for most of my career. Without some sense of how often operations beat placebo when properly tested, there is no calibrated prior to read a new trial against.
2
A Controlled Trial of Arthroscopic Surgery for Osteoarthritis of the Knee
NEJM · 2002 · Moseley JB Jr, O'Malley K, Petersen NJ et al. · Sham-controlled RCT, n=180
What it does: Randomises patients with knee osteoarthritis to arthroscopic débridement, arthroscopic lavage, or placebo skin incisions with no arthroscope inserted. At no point across 24 months did either real operation produce less pain or better function than the placebo. The confidence intervals excluded any clinically meaningful difference.
Epistemological angle: Before this trial, the evidence for knee arthroscopy in osteoarthritis was consistent, abundant, and entirely uncontrolled. Patients did improve. The inference from "they improve after surgery" to "the surgery improved them" is the one this design breaks, and it turns out to have been carrying the whole weight. Regression to the mean, natural history, and the substantial expectation effects of an operation account for the observed benefit without requiring the mechanism anyone believed in.
My take: The uncomfortable part is not that the operation failed. It is that thousands of surgeons had watched patients get better and drawn the obvious conclusion, and the obvious conclusion was wrong. Clinical experience is a real form of knowledge, but it is structurally blind to exactly this failure mode, and no amount of it accumulates into a control group.
N Engl J Med 2002;347(2):81–88 · PMID 12110735 ·
doi:10.1056/NEJMoa013259 · editorial: Felson & Buckwalter, NEJM 2002;347(2):132–133, PMID 12110742
3
Arthroscopic Partial Meniscectomy versus Sham Surgery for a Degenerative Meniscal Tear
NEJM · 2013 · Sihvonen R, Paavola M, Malmivaara A et al. (FIDELITY) · Double-blind sham-controlled RCT, n=146
What it does: Takes patients with a degenerative medial meniscal tear and no osteoarthritis — the classic indication — and randomises them to partial meniscectomy or a sham arthroscopy. At twelve months there was no difference on Lysholm, WOMET, or pain after exercise.
Epistemological angle: The tear was the reference class. It was the finding that assigned the patient to the operable group, that justified the arthroscope, and that explained the pain afterwards. FIDELITY shows the assignment does not license the operation: removing the thing did not remove the symptom. That is a weaker claim than showing the tear to be causally inert, and the difference is worth holding onto — a process may be irreversible by the time it is found, or the surgery's own costs may offset what it gains. But it is enough for the clinical question, because the assignment's entire purpose was to justify removal. Read alongside entry
12, which establishes that most people of this age have such a tear and most of them have no pain, the two papers form a complete argument — here is the base rate that should have warned us, and here is the trial that eventually did.
My take: The response is as instructive as the result. A formal rebuttal appeared in Arthroscopy within months arguing the trial was flawed and unrepresentative, and the practice pattern took the better part of a decade to move. That is not corruption. That is what it looks like when disconfirming evidence arrives about a procedure a specialty has organised itself around.
N Engl J Med 2013;369(26):2515–2524 · PMID 24369076 ·
doi:10.1056/NEJMoa1305189 · rebuttal: Elattrache N, Lattermann C, Hannon M, Cole B. Arthroscopy 2014;30(5):542–543, PMID 24642105
4
The Vertebroplasty Sequence: Two Null Trials, Then One That Was Not
NEJM · 2009 · Buchbinder R et al. (n=78) and Kallmes DF et al., INVEST (n=131), independent sham-controlled RCTs published in the same issue · and Lancet · 2016 · Clark W, Bird P, Gonski P et al., VAPOUR, double-blind placebo-controlled RCT (n=120)
What it does: Two research groups on opposite sides of the world independently randomised patients with painful osteoporotic vertebral fractures to cement injection or a sham procedure. Neither found any advantage at any time point; both arms improved substantially. Seven years later VAPOUR repeated the design in a deliberately different population — fractures under six weeks old, severe pain, a specific filling technique — and vertebroplasty won, with 44% against 21% below a pain score of 4 out of 10 at fourteen days.
Epistemological angle: Two lessons in sequence, and the second undoes the comfortable reading of the first. A single negative trial invites the reply that this trial was flawed; two independent trials in different countries with different designs and the same null move the question from the trials onto the intervention, and the epistemic gain comes from the independence rather than from the combined sample size. INVEST adds that blinding and crossover-resistance are separate problems, with 51% of its sham arm crossing to real vertebroplasty by three months against 13% the other way. Then VAPOUR states the reference class problem as a trial. It did not refute the 2009 results; it redrew the boundary of who was being asked about. "Does vertebroplasty work" was never a well-formed question — only "in fractures of what age, what severity, filled to what extent" is answerable, and every null result carries an implicit population that the headline discards. Whether VAPOUR's chosen class is the right one remains contested, largely on the adequacy of its blinding.
My take: Cement stabilises a fracture. It is difficult to think of a more mechanically obvious intervention in all of orthopaedics, and in two trials it did nothing. I would hold onto that whenever I catch myself reasoning from mechanism to benefit, which is most days. But VAPOUR is why this is one entry rather than two, and it is the more useful half. It is easy to enjoy sham trials when they confirm a suspicion that we operate too much. The discipline is in accepting the same design when it says the opposite, and in noticing that both answers are about populations rather than about procedures.
5
The Two Decompression Trials
Lancet · 2018 · Beard DJ, Rees JL, Cook JA et al. (CSAW, n=313) · and BMJ · 2018 · Paavola M, Malmivaara A, Taimela S et al. (FIMPACT, n=210) · Independent placebo-controlled RCTs
What it does: Both trials test arthroscopic subacromial decompression against a placebo arthroscopy in which the arthroscope is introduced and the bone and soft tissue are left alone. CSAW ran three arms — decompression, arthroscopy only, no treatment — across 32 UK hospitals and 51 surgeons, and found no difference between the two surgical groups on the Oxford Shoulder Score at six months. Both surgical arms beat no treatment by a margin the investigators judged not clinically important. FIMPACT ran the comparison in Finland with pain at rest and on activity at 24 months as co-primary outcomes and reached the same conclusion: decompression offered no benefit over diagnostic arthroscopy.
Epistemological angle: The independence argument of entry
4, repeated in a second procedure, and with a sharper design. CSAW's three-arm structure separates two questions that a two-arm sham trial has to answer together: what does the specific surgical act contribute, and what does the whole surgical episode contribute. The answer is that the act contributes nothing detectable and the episode contributes a little. That decomposition is what makes these trials the natural terminus of the arc begun at entry
8: Neer's mechanism was proposed in 1972, the diagnosis built on it, the operation built on the diagnosis, and forty-six years later the mechanism-specific step turned out to be the part that was doing nothing.
My take: Promoting these out of a footnote is a correction to an earlier version of this list, which cited them only in passing under Neer. They deserve to stand on their own, and the reason is CSAW's middle arm. A two-arm sham trial tells you the operation is no better than nothing much; a three-arm trial tells you which part of it was inert. We should be designing more trials that can answer that question, and I do not know why we design so few.
6
Is the Placebo Powerless? An Analysis of Clinical Trials Comparing Placebo with No Treatment
NEJM · 2001 · Hróbjartsson A, Gøtzsche PC · Systematic review, 130 trials
What it does: Reviews trials containing both a placebo arm and an untreated arm — the only design that can measure the placebo effect itself — and finds little evidence of a large general clinical effect, with small effects on subjective outcomes such as pain. Their 2010 Cochrane update adds that physical placebos, including sham procedures, produce larger responses than pills.
Epistemological angle: This complicates the lazy reading of every trial above. "It was just the placebo effect" sounds like an explanation, but if placebo is generally weak, then a sham arm matching a real operation is a far harsher verdict than that phrase implies — it means the operation was doing very little, not that placebo was doing a great deal. And the finding that procedural placebos outperform pill placebos is what makes surgical sham controls necessary rather than borrowable: surgery could not import its intuitions from drug trials, because the placebo in surgery is a different and larger thing.
My take: This is the entry most likely to be skipped and it is the one that makes the rest rigorous. Without it, "sham was as good as surgery" gets absorbed as a comforting story about the power of the mind, when the harder reading is that the operation was close to inert.
N Engl J Med 2001;344(21):1594–1602 · PMID 11372012 ·
doi:10.1056/NEJM200105243442106 · update: Cochrane Database Syst Rev 2010;(1):CD003974, PMID 20091554
7
Vertebral Augmentation After Recent Randomized Controlled Trials: A New Rise in Kyphoplasty Volumes
JACR · 2015 · Cox M, Levin DC, Parker L, Morrison W, Long S, Rao VM · Medicare billing analysis, 2006–2013
What it does: Tracks procedure volumes after the 2009 sham-controlled vertebroplasty trials. Vertebroplasty fell sharply. Kyphoplasty — a mechanistically similar procedure never subjected to the same test — dipped and then rose, partly offsetting the decline in total vertebral augmentation.
Epistemological angle: The cleanest empirical demonstration of the adopt-versus-abandon asymmetry available in orthopaedics. Kyphoplasty was never held to the standard eventually applied to vertebroplasty; it entered practice on plausibility and stayed there. When disconfirming evidence arrived for the tested procedure, demand migrated to the untested cousin. No individual made an irrational decision, and the aggregate result was that a sham-controlled refutation produced a change in billing codes rather than a change in whether patients had cement injected.
My take: This is what makes me sceptical that better evidence alone fixes anything. The evidence arrived, it was excellent, it was widely publicised, and the practice found a route around it that required nobody to disagree with it.
Also worth reading
Macklin R versus Freeman TB, Vawter DE, Leaverton PE et al. Is sham surgery permissible? Opposing essays, N Engl J Med 1999;341(13):992–996, PMID 10498498, doi:10.1056/NEJM199909233411312 · and 341(13):988–992, PMID 10498497, doi:10.1056/NEJM199909233411311. The ethical cost of the design that makes this theme possible, argued by two competent parties neither of whom is wrong. Developed further in Horng S, Miller FG, N Engl J Med 2002;347(2):137–139, PMID 12110744, doi:10.1056/NEJMsb010576.
Theme Two
What a name does
Does the label describe the disease, or help bring it about?
A diagnosis is supposed to be a report on the world. In musculoskeletal medicine it is frequently something else: an assignment to a class, made on the basis of a finding whose relationship to the symptom was never established, which then shapes what the patient believes, what the surgeon offers, and what happens next. These three entries track a label from its invention to the randomised experiments showing what it does once spoken. The prior question — whether the named entity corresponds to anything the world divides that way — is theme ten's business.
8
Anterior Acromioplasty for the Chronic Impingement Syndrome in the Shoulder: A Preliminary Report
JBJS Am · 1972 · Neer CS 2nd · Uncontrolled case series
What it does: Proposes that chronic shoulder pain arises from mechanical impingement of the rotator cuff beneath the anterior acromion, and describes an operation to relieve it by removing the offending bone. There is no control group, no blinding, no comparator. Neer called it preliminary in the title.
Epistemological angle: This is the origin of a diagnostic category, an operation, and eventually an industry. A mechanism was proposed to explain an observation; a name was attached to the mechanism; and the name was thereafter used as though the mechanism had been demonstrated. Forty-six years later, the trials at entry
5 found that subacromial decompression does not outperform a placebo arthroscopy, and by then Lewis had already argued that the label encoded a causal story nobody had tested. The arc runs from hypothesis to diagnosis to industry to disconfirmation across half a century, and at no point in the chain did anyone go back and examine the first step.
My take: Neer wrote "preliminary" in the title and meant it. What happened afterwards happened downstream of him: a hypothesis became a diagnosis, the diagnosis licensed an operation, and the premise went unrechecked for forty-five years.
J Bone Joint Surg Am 1972;54(1):41–50 · PMID 5054450 · no DOI; pre-abstract era
9
Increased Absenteeism from Work After Detection and Labeling of Hypertensive Patients
NEJM · 1978 · Haynes RB, Sackett DL, Taylor DW, Gibson ES, Johnson AL · Industrial screening cohort with randomised compliance intervention
What it does: Screens steelworkers for hypertension and tracks absenteeism afterwards. Among those previously unaware they were hypertensive, absenteeism rose by 5.2 days per year — an 80% increase against a 9% rise in the general workforce over the same period. The rise was associated with becoming aware of the condition, and was unrelated to the severity of the hypertension, to whether treatment was started, or to whether blood pressure was subsequently controlled.
Epistemological angle: Everything the modern labelling trial at entry
10 demonstrates was shown here, in a different disease, forty-odd years earlier. The design isolates the label unusually cleanly: the physiological state was present before screening and unchanged by it, so the only new object in the world is the patient's knowledge of a category they have been placed in. That the effect was independent of severity and of treatment is what makes it an epistemological finding rather than a clinical one — the number on the sphygmomanometer did not predict the harm, the act of naming did. This is the entry that establishes the phenomenon is a property of diagnosis as such, not a quirk of musculoskeletal language.
My take: Sackett co-authored this, which is worth noticing given what he went on to found. The man most associated with evidence-based medicine published early evidence that the diagnostic act carries a cost the evidence base does not measure. I would not have predicted that, and I think it complicates the caricature of EBM as naively realist about diagnosis.
10
Diagnostic Labels for Rotator Cuff Disease Can Increase People's Perceived Need for Shoulder Surgery
JOSPT · 2021 · Zadro JR, O'Keeffe M, Ferreira GE, Haas R, Harris IA, Buchbinder R, Maher CG · Six-arm randomised trial, n=1308
What it does: Presents an identical clinical scenario to 1,308 people, randomising only the diagnostic label attached to it. Calling the presentation a "rotator cuff tear" rather than "bursitis" significantly raised the perceived need for surgery and for imaging.
Epistemological angle: The clinical content was held constant. The label added no information about the shoulder, and it changed behaviour anyway. That makes it an input rather than only a report — at minimum an input into what the patient wants done, and the interview study in this theme’s further reading gives reason to think it runs further than that. In reference-class terms the label is the assignment mechanism, and here the assignment is doing work that no fact about the joint supports. What the design shows directly is an effect on stated intentions; the step from intentions to the course of the illness is an inference this study invites rather than establishes.
My take: We suspected this. What we lacked was a randomised design with the clinical facts pinned down, which is what makes this the entry rather than the accumulated clinical impression. It seems to me a reason to think carefully about how we word reports, though I recognise that is easier said than done.
Also worth reading
O'Keeffe M, Ferreira GE, Harris IA, Darlow B, Buchbinder R et al. Effect of diagnostic labelling on management intentions for non-specific low back pain. Eur J Pain 2022;26(7):1532–1545 · PMID 35616226 · doi:10.1002/ejp.1981. Replicates entry 10 in the spine, with the effect largest in people already most exposed to overtreatment.
Darlow B, Dowell A, Baxter GD, Mathieson F, Perry M, Dean S. The enduring impact of what clinicians say to people with low back pain. Ann Fam Med 2013;11(6):527–534 · PMID 24218376 · doi:10.1370/afm.1518. Twenty-three interviews showing what a label does over years rather than what it does to a stated intention. Supplies the mechanism the vignette trials can only infer.
Theme Three
The class nobody checked
How common is this finding in people who feel fine?
The simplest question in this collection, and the one most reliably skipped. Before a finding can be used to explain a symptom, someone has to know how often it turns up in people with no symptom at all. Spine, knee and hip here; shoulder and the pooled age-gradient in the further reading. Four decades, and the same answer every time.
11
Magnetic Resonance Imaging of the Lumbar Spine in People without Back Pain
NEJM · 1994 · Jensen MC, Brant-Zawadzki MN, Obuchowski N, Modic MT, Malkasian D, Ross JS · Blinded cross-sectional study, n=98
What it does: Scans people with no back pain, mixes their images among symptomatic ones, and has neuroradiologists read them blind. 52% had a bulging disc, 27% a protrusion. Only 36% had normal discs at every level.
Epistemological angle: The founding document of the theme. A positive finding and a diagnosis are different objects, and the difference is precisely the base rate that nobody had measured. The blinding matters more than it looks: the radiologists could not know which images came from which group, which forecloses the objection that the abnormalities were read into the scans. What the study establishes is not that imaging is useless but that its output is uninterpretable without a denominator.
My take: Thirty years old, and practice has been slow to shift. The prevalence figure is the lesser lesson; the greater one is how much more than a number it takes to change what we do.
12
Incidental Meniscal Findings on Knee MRI in Middle-Aged and Elderly Persons
NEJM · 2008 · Englund M, Guermazi A, Gale D, Hunter DJ, Aliabadi P, Clancy M, Felson DT · Framingham population cohort, n=991
What it does: Uses a population sample rather than volunteers — people were not selected for having or lacking knee symptoms. Meniscal tear prevalence ran from 19% in women aged 50 to 59 up to 56% in men aged 70 to 90. Of those with a tear, 61% reported no knee pain, aching or stiffness in the previous month.
Epistemological angle: The population sampling is what makes this the strongest entry in the theme. Volunteer studies invite the objection that volunteers are unrepresentative; a community cohort does not. And the number that matters is the 61%, because it converts the abstract point about base rates into a direct statement about diagnostic value: most people carrying this finding are not in pain, so its presence in a patient who is in pain cannot on its own explain the pain. That is a claim about what the finding licenses, not about what it does. A tear may well be contributing something in some people — high prevalence in the well does not refute causation in the ill, any more than the fact that most smokers never develop lung cancer refutes the effect of smoking. What high prevalence destroys is the finding's ability to tell you which people.
My take: Published five years before FIDELITY, in the same journal, by people the field respected. Much of what would have predicted that trial's result was already in print. The evidence was not missing so much as unpersuasive at the time, which is a harder problem than ignorance and not obviously anyone's failing.
13
Prevalence of Femoroacetabular Impingement Imaging Findings in Asymptomatic Volunteers
Arthroscopy · 2015 · Frank JM, Harris JD, Erickson BJ, Slikker W, Bush-Joseph CA, Salata MJ, Nho SJ · Systematic review, 26 studies, 2,114 asymptomatic hips
What it does: Pools imaging studies of hips in people with no hip symptoms, mean age 25. Cam morphology was present in 37% overall — 54.8% in athletes against 23.1% in the general population. Pincer morphology was present in 67%, though the authors note it was poorly and inconsistently defined across studies. Of the seven studies reporting on the labrum, labral injury was found on MRI in 68.1% of asymptomatic hips.
Epistemological angle: The theme's fourth joint, and the one where the lesson is currently live rather than historical. Spine, shoulder and knee learned this in 1994, 1995 and 2008; hip arthroscopy volumes grew through the 2010s on findings whose asymptomatic prevalence was being established at the same time. The pincer figure carries a second lesson, one the authors state plainly: a prevalence estimate inherits the definition used to generate it, and when only four of twenty-six studies defined the deformity explicitly, 67% is a number about the literature as much as about hips. That is a different failure from the one at entries
11 and
12 — there the denominator was unmeasured, here it is measured against a moving numerator.
My take: Two-thirds of pain-free young hips showing labral injury is the same shape of number as Sher's shoulders, arriving twenty years later in a field that had the earlier examples available to it. I would like to think that means we act on it faster this time. The volumes suggest otherwise, and I do not exempt myself from the reasons why.
Also worth reading
Sher JS, Uribe JW, Posada A, Murphy BJ, Zlatkin MB. Abnormal findings on magnetic resonance images of asymptomatic shoulders. J Bone Joint Surg Am 1995;77(1):10–15 · PMID 7822341 · doi:10.2106/00004623-199501000-00002. 34% of pain-free shoulders with a cuff tear, 54% over the age of sixty. Makes the point in the tissue surgeons are most confident about.
Brinjikji W, Luetmer PH, Comstock B, Deyo RA, Jarvik JG et al. Systematic literature review of imaging features of spinal degeneration in asymptomatic populations. AJNR Am J Neuroradiol 2015;36(4):811–816 · PMID 25430861 · doi:10.3174/ajnr.A4173. Establishes that the base rate is a function of age rather than a number: disc degeneration in pain-free people from 37% at twenty to 96% at eighty.
Boden SD, Davis DO, Dina TS, Patronas NJ, Wiesel SW. Abnormal magnetic-resonance scans of the lumbar spine in asymptomatic subjects. J Bone Joint Surg Am 1990;72(3):403–408 · PMID 2312537 · no DOI in record. Anticipates entry 11 by four years.
Theme Four
Bayes, the prior, and the reference class
What did the p-value throw away, and what would a prior have kept?
Theme three establishes that base rates exist and are often extreme. This theme is about what happens next. It opens with four papers that state the problem in general form — the base rate of true findings in a literature, how often established practice reverses when tested, the question a p-value does not answer, and the identity between a prior and a reference class — and then turns to whether clinicians in fact reason this way. Two entries measure surgeons directly, and the results are not flattering to any of us.
14
Why Most Published Research Findings Are False
PLoS Med · 2005 · Ioannidis JPA · Analytical essay with simulations · general science
What it does: Models the probability that a published research claim is true as a function of study power, bias, the number of teams working the same question, and — the load-bearing term — the ratio of true to null relationships among those a field chooses to test. Under most realistic combinations of these, the post-study probability that a claimed positive finding is true falls below one half.
Epistemological angle: This is theme three's argument turned on the literature rather than the patient. A published positive finding is a test result. Interpreting it requires a prior, and the prior is the proportion of hypotheses in that field that are true — a quantity nobody measures and everybody implicitly sets at something close to one. Ioannidis makes the reference class of a research claim explicit: it belongs to the set of studies of its design, in its field, at its power, and its credibility is a property of that set before it is a property of the paper. His last observation is the sharpest, and it is a claim about theme nine as much as this one: where prejudice and financial interest are strong, a published finding may be an accurate measurement of the field's prevailing bias rather than of the world.
My take: The title is doing rhetorical work and the paper is more careful than the title. It is a conditional argument, not a census, and how much of it applies to orthopaedic surgical trials depends on parameters we have never estimated for ourselves. I would rather read it as a framework we have not yet run our own numbers through than as a verdict already delivered on us. That said, the exercise of guessing at those parameters for our specialty is uncomfortable in a way I think is instructive.
15
The Reversal Rate: How Often Established Practice Fails When Tested
Mayo Clin Proc · 2013 · Prasad V, Vandross A, Toomey C et al. · 2,044 original articles over ten years · and JAMA · 2005 · Ioannidis JPA · 49 highly cited clinical studies
What it does: Prasad classifies every original article published over a decade in one high-impact journal by what it did to practice. Of 363 articles testing an established practice, 146 (40.2%) reversed it and 138 (38.0%) reaffirmed it. Ioannidis asks the same question of celebrated findings rather than of standard care: of 49 highly cited clinical studies, 45 claimed an intervention effective, and of those, seven were later contradicted outright and seven more had found effects stronger than subsequent work. Twenty replicated. The breakdown by design is the sharp part — five of six highly cited non-randomised studies were contradicted or attenuated, against nine of 39 randomised trials (p=0.008), and among the trials the vulnerable ones were the small ones.
Epistemological angle: Together these supply the prior that entry
14 models theoretically. When an established practice is finally subjected to a proper test it fails roughly as often as it survives, and a celebrated finding has about a one-in-three chance of shrinking or reversing. That number should govern how much evidential weight "this is standard of care" and "this is the landmark trial" carry, and in practice both carry far more. Prasad's denominators add a second finding that is easy to miss: 73% of the literature examined new practices and 27% examined existing ones. The field spends roughly three times the effort finding things to start as finding things to stop, and then treats the untested incumbent as the safe default.
My take: Between them these are the most useful numbers in the collection for a surgeon's own calibration, and both sit in general medicine rather than in ours. Nobody has run either analysis on the orthopaedic literature. That would be a straightforward and worthwhile study, and my guess is that our rate is not better. I would rather be shown wrong by data than continue guessing.
16
Toward Evidence-Based Medical Statistics: The P Value Fallacy and The Bayes Factor
Ann Intern Med · 1999 · Goodman SN · Two-part methodological essay · general medicine
What it does: Part 1 argues that the standard hypothesis-testing apparatus is an amalgam of two incompatible statistical philosophies, and that the p-value fallacy is the belief that a single number can carry both the long-run error behaviour of a procedure and the evidential meaning of the particular result in front of you. Part 2 proposes the Bayes factor as the measure that separates them, and shows that p-values systematically overstate the evidence against the null.
Epistemological angle: Most of this collection asserts at some point that the p-value answers a question the surgeon did not ask. This is where that claim is actually argued, and the argument is more interesting than the slogan. Goodman's point is not that frequentist statistics are wrong but that they are answering a question about procedures — how often would this method mislead me in the long run — and that clinicians read the answer as though it were about this result, this patient population, this decision. The Bayes factor is offered as the piece that makes the difference visible: evidential strength first, and then, separately and explicitly, whatever background knowledge you bring to it. Read this before entry
21, which is the same argument carried out on real orthopaedic data.
My take: Part 1 is the one to read if you read only one, and it is twenty-five years old, which should give us pause. The critique was available, clear, published in a journal every academic clinician reads, and the field's statistical practice is essentially unchanged. That is a fact about incentives and training rather than about the argument's quality, and it is the kind of fact this collection keeps running into.
17
The Reference Class Problem Is Your Problem Too
Synthese · 2007 · Hájek A · Philosophical paper · outside PubMed's scope; citation verified against the journal record
What it does: Argues that the reference class problem — that any individual belongs to indefinitely many groups with different frequencies, and nothing in the evidence selects among them — is not a local defect of frequentism to be escaped by adopting some other interpretation of probability. Subjectivism, logical probability and propensity accounts each face a structurally identical problem at some point in their machinery. For the Bayesian, it surfaces in the choice of prior.
Epistemological angle: This is the beam the whole collection rests on, and it belonged in the list from the start rather than only in the introduction. Every entry in theme three is a reference class problem; so is the disagreement in entry
24, the threshold in entry
25, and the population question in entry
4. Hájek's contribution is to close the exit. The comfortable reading of theme four is that if surgeons would only adopt priors, the arbitrariness would go away. It does not go away; it moves. What Bayesian practice buys is not the elimination of the choice but its visibility — the prior has to be written down, which means it can be argued with, varied, and shown to matter. That is a real gain, and it is a smaller one than it is usually sold as. The point returns in mechanical form at entry
42: a trained model's reference class is whatever its training data made salient, the model does not choose it and cannot report it, and the question of which class a system was right about turns out to be the only one that governs whether its accuracy transfers to your patient.
My take: I find this the most clarifying paper on the list and the least likely to be read from it. The reason to include it anyway is that without it the Bayesian material reads as a technical fix for a technical problem, when the actual situation is that the judgment cannot be eliminated from anywhere and the best we can do is put it where people can see it. That is a modest conclusion and I think it is the true one.
18
Base-Rate Neglect Across Thirty-Six Years
NEJM 1978, Casscells W, Schoenberger A, Graboys TB · and JAMA Intern Med 2014, Manrai AK, Bhatia G, Strymish J, Kohane IS, Jain SH · general medicine
What it does: Casscells posed a problem to physicians, residents and students: a disease with prevalence one in a thousand, a test with a 5% false-positive rate, a positive result — what is the probability the patient has the disease? The modal answer was 95%. The answer is about 5%. Manrai re-ran it with modern clinicians thirty-six years later and got the same failure.
Epistemological angle: Taught as a pair because the replication is the finding. A single 1978 result invites the reply that training has improved since. It has not. The error is stable across four decades of curriculum reform, which suggests it is not a gap in what clinicians were taught but a feature of how the reasoning system works under this particular load. What gets neglected is the denominator, and the denominator is the reference class.
My take: The specific shape of the error is worth noting: people substitute the test's false-positive rate for the probability of disease. They are not calculating badly, they are answering a different and easier question without noticing the swap.
19
Cognitive Biases in Orthopaedic Surgery
JAAOS · 2021 · Janssen SJ, Teunis T, Ring D, Parisien R · Vignette survey, n=196 orthopaedic surgeons
What it does: Puts the Casscells problem and several others directly to orthopaedic surgeons. Across three base-rate vignettes, 43%, 88% and 35% chose answers consistent with base-rate neglect. Confirmation bias appeared in 11% to 51% depending on the scenario. Randomising the anchor moved responses by 22 percentage points; changing the frame moved them by 16.
Epistemological angle: This is the entry that brings the general literature home. Everything in entry
18 was demonstrated in physicians broadly; this demonstrates it in our specialty, with our vignettes, at rates that make it impossible to treat as somebody else's problem. The anchoring and framing results are arguably worse than the base-rate ones, because they show the same surgeon gives different answers to the same clinical question depending on an arbitrary number they were shown first.
My take: The 88% vignette gave me pause. This is not a minority struggling with a hard question but most of a thoughtful sample, which suggests the question is genuinely hard. The framing effect adds something worth sitting with: some part of surgeon-to-surgeon variation may reflect different anchors rather than substantive disagreement.
20
Musculoskeletal Surgeons Use Mixed Reasoning Rather than Pure Bayesian Strategies in Clinical Practice
PLoS One · 2026 · Parisien R, Drost A, Razi A, Ramtin S, Ring D, Janssen SJ · Survey, n=153 Science of Variation Group
What it does: Presents surgeons with eight scenarios involving test and treatment decisions under uncertainty, scoring each response on a four-point scale from non-Bayesian to fully Bayesian. Mean score 3.0. Fully non-Bayesian reasoning was rare at 8.6%; fully Bayesian reasoning accounted for 29%. Most responses were mixed — prior probability acknowledged but underweighted, without explicit updating. 85% of surgeons reasoned fully Bayesianly at least once; 42% reasoned non-Bayesianly at least once.
Epistemological angle: The interesting result is the inconsistency, not the average. The same surgeon reasons well in one scenario and badly in another, which means Bayesian competence is not a stable trait being measured. The Cronbach alpha of 0.43 makes this explicit and the authors say so: the scenarios were not tapping one underlying construct. That is a more useful finding than a clean score would have been, because it relocates the problem from the person to the situation — and situations can be redesigned in a way that dispositions cannot.
My take: A low alpha is usually reported apologetically. Here it is the point: "does this surgeon reason Bayesianly" turns out not to be a well-formed question, which is itself an argument for decision aids over exhortation. One limitation cuts in a particular direction: the Science of Variation Group is self-selected and unusually reflective, so the wider picture may well be less favourable.
21
High-Dose Dual Antibiotic Cement for Hip Hemiarthroplasty: A Post-Hoc Bayesian Analysis
Bone & Joint Journal · 2025 · Farrow L, Hudson J, George A, Reed MR, Campbell MK · Bayesian reanalysis of the WHiTE 8 RCT, n=4,406
What it does: Takes a large randomised trial whose frequentist analysis was null — odds ratio 1.44, 95% confidence interval 0.88 to 2.37, filed as no benefit — and reanalyses it with explicit priors. The posterior probability of at least some benefit from dual-antibiotic cement was 92% to 98%, depending on which prior was used.
Epistemological angle: The best demonstration in the orthopaedic literature of what the frequentist convention discards, precisely because the contrast is within a single dataset rather than between studies. The p-value answered whether the data were surprising under a null nobody believed. The posterior answers how likely the treatment is to help, which is the question a surgeon actually has. Note also that the answer depends on the prior — that is not a weakness but the honest exposure of an assumption that a frequentist analysis makes silently, and entry
17 is the reason it cannot be exposed any further than that.
My take: "No significant difference" was never a finding about the world; it was a statement about a threshold. A trial of four thousand patients told us something quite specific about infection risk and the conventional analysis reported it as nothing.
22
Does the Surgical Innovation Evidence Apply to You? How Surgeon Volume and Experience Determine Whether Adoption Helps or Harms
Clin Orthop Relat Res · 2026 · Parisien R · Volume-stratified break-even decision model
What it does: Builds a decision model for whether adopting a new surgical approach helps or harms a given surgeon, using direct anterior versus posterior hemiarthroplasty as the worked case. Three parameters carry it: the number needed to treat from the trial evidence, a learning-curve decay constant governing how quickly early excess harm resolves, and a weight expressing how much a harm counts against a benefit. Break-even is reached at different annual volumes under different assumptions, and below that volume adoption is expected to harm. An interactive version lets the reader enter their own figures.
Epistemological angle: Nearly all discussion of the reference class problem in medicine concerns the patient — which group does this person belong to, and does the trial's population contain them. This applies the same problem to the operator, which is an unusual move and, in surgery, the more consequential one. A trial's effect estimate is an average over that trial's surgeons, who were selected for competence with the technique under test; the surgeon deciding whether to adopt it is very often not a member of that class, and no feature of the published result tells them so. Where entry
37 argues that the instrument's variability should be built into allocation, this argues that it should be built into the reading of a result already published, and makes the class membership a quantity rather than a caveat. It is also the only entry here where the reference-class choice is operationalised rather than described: Hájek's point at entry
17 is that the choice cannot be eliminated, only made visible, and a model with the class as an explicit parameter is one way of making it visible enough to argue with.
My take: I wrote this, so the criticisms should come from me. The decay constant and the harm weight are the whole model, and neither has been measured — they are elicited, and a conclusion sensitive to two unmeasured parameters is an argument structure rather than a finding. It is illustrated on a single procedure pair, which is narrower than the claim it is making. And the deeper problem is one the paper concedes too quietly: it individuates surgeons by annual volume because volume is what registries record, not because volume is the variable that carves. Case mix, assistant quality and institutional support may matter more, and if they do, the reference class problem has simply reappeared inside the proposed solution. That is Hájek's point arriving at my own door, and I do not think a model can escape it — only make the choice legible and let the reader disagree with it. What I would defend is the direction: we have spent a great deal of effort asking whether a patient resembles a trial population and almost none asking whether the surgeon does.
Clin Orthop Relat Res 2026 · in press; volume, pages, PMID and DOI to be supplied at publication · interactive model at
rcpmodel.netlify.app
Also worth reading
Gill CJ, Sabin L, Schmid CH. Why clinicians are natural Bayesians. BMJ 2005;330(7499):1080–1083 · PMID 15879401 · doi:10.1136/bmj.330.7499.1080. Worth reading before the survey entries because it stops them landing as an insult. Nobody reasons from a blank prior; the question is whether the prior in use is the one you would endorse if you looked at it.
The Thessaly test, rise and collapse. Karachalios T et al., J Bone Joint Surg Am 2005;87(5):955–962, PMID 15866956, doi:10.2106/JBJS.D.02338 · Konan S et al., Knee Surg Sports Traumatol Arthrosc 2009;17(7):806–811, PMID 19399477, doi:10.1007/s00167-009-0803-3 · Goossens P et al., J Orthop Sports Phys Ther 2015;45(1):18–24, PMID 25420009, doi:10.2519/jospt.2015.5215 · Blyth M et al., Health Technol Assess 2015;19(62):1–62, PMID 26243431, doi:10.3310/hta19620. A single-centre accuracy claim of 94–96% collapsing to 54–59% as replication became independent — entry 15's general finding as one specific story. Demoted only because its base-rate lesson duplicates theme three; it remains the best teaching arc in the collection and belongs in any course built from this list.
Theme Five
Proxies, surrogates, and moving the goalposts
Who decided what counts as success, and what did they have at stake?
Every surrogate endpoint is a claim that one reference class stands in for another — that radiographic union means a usable limb, that asymptomatic clot means fatal embolism, that a score above some threshold means a patient who is glad they had the operation. The formal condition under which such a claim holds is that the intervention's effect on the real outcome be fully mediated by the surrogate; correlation between the two is not enough and never was. The entries here are about what happens when nobody checks, and about the special case where a field sets its own passing mark.
23
The Cardiac Arrhythmia Suppression Trial (CAST)
NEJM · 1989 and 1991 · CAST Investigators; Echt DS, Liebson PR, Mitchell LB et al. · Placebo-controlled RCT, n=1,498 in the encainide and flecainide arms
What it does: Tests the hypothesis that suppressing ventricular ectopy after myocardial infarction reduces sudden death. Patients whose ectopy was successfully suppressed by encainide or flecainide were randomised to continue the drug or to placebo. The encainide and flecainide arms were stopped early: 43 arrhythmic deaths on drug against 16 on placebo, and total mortality 7.7% against 3.0% in the preliminary report. The drugs did exactly what they were designed to do to the surrogate, and killed people.
Epistemological angle: The example the surrogate literature is built on, given its own entry here because the structure deserves examination rather than citation. Every element of the reasoning was sound in isolation. Ectopy predicts sudden death — true, and well established. These drugs suppress ectopy — true, and confirmed in every patient before randomisation. Therefore suppressing ectopy should prevent death — false, and false in a way no amount of further observational data would have exposed. The error is the assumption that a marker on the causal pathway is a lever on it: ectopy was a symptom of a diseased myocardium, and the drugs suppressed the symptom while making the myocardium worse. This is the surrogate fallacy at maximum strength, with the additional detail that the trial design itself enrolled only responders, which should have made the result more favourable and did not.
My take: I keep this one close because the reasoning is not stupid. It is the reasoning I use. Restore the anatomy and function follows (entry
27), achieve radiographic union and the limb works (entry
50), reattach the tendon and the shoulder recovers (entry
39) — these are the same argument in a different tissue, and cardiology ran the experiment we have mostly not run. The comfort is that this happened in a field with better trials than ours. The discomfort is the same fact.
24
AAOS and ACCP Guidelines for Venous Thromboembolism Prevention Differ: What Are the Implications?
Chest · 2009 · Eikelboom JW, Karthikeyan G, Fagel N, Hirsh J · Comparative guideline critique
What it does: Documents that two expert bodies issued directly conflicting recommendations for thromboprophylaxis after hip and knee arthroplasty from the same trial evidence. The American College of Chest Physicians accepted venographic deep vein thrombosis — largely asymptomatic clot detected on imaging — as a valid efficacy outcome. The AAOS rejected it as an unproven proxy for pulmonary embolism.
Epistemological angle: The clearest demonstration available that evidence underdetermines practice. Neither body was incompetent, neither had different data, and their advice to surgeons was opposite. The entire disagreement reduces to whether one reference class — asymptomatic venographic clot — licenses inference to another, fatal embolism. That is a question about surrogate validity, not about the trials, and no additional trial of the same design could have settled it.
My take: When guidelines conflict, the instinct is to ask which committee was captured or careless. Usually neither. They disagreed about what would count as an answer, which is a deeper disagreement than any about the data and one that more data does not touch.
25
Hemiarthroplasty for the Rotator Cuff-Deficient Shoulder
JBJS Am · 2008 · Goldberg SS, Bell JE, Kim HJ, Bak SF, Levine WN, Bigliani LU · Retrospective case series, 34 shoulders
What it does: Reports that 26 of 34 shoulders, or 76%, "satisfied the limited goals criteria described by Neer et al." Forward elevation improved from 78 to 111 degrees, external rotation from 15 to 38. The success threshold was supplied by the surgical tradition that developed the operation.
Epistemological angle: A threshold calibrated to what a procedure can achieve cannot then be used as evidence that it achieves anything. The criterion is not literally unfailable — eight of the thirty-four shoulders did fail it, so the study could have come out worse than it did. What it lacks is any anchor outside the tradition being assessed: nothing but that tradition determines where the bar sits, so a poor result and a good one are both measured against a standard the operation's own advocates set. And the pattern is not isolated: Williams and Rockwood reported 86% satisfactory by this standard in 1996, Sanchez-Sotelo and Cofield 67% in 2001. Three decades of apparently comparable evidence, all resting on a benchmark with no external referent.
My take: Normal function genuinely is the wrong benchmark for a cuff-deficient shoulder, and the alternative genuinely should come from someone who understands the pathology. The loop just closed, and the field cited it to itself for thirty years without noticing that no patient had ever been asked what threshold mattered to them.
J Bone Joint Surg Am 2008;90(3):554–559 · PMID 18310705 ·
doi:10.2106/JBJS.F.01029 · concept originates with Neer CS 2nd, Craig EV, Fukuda H, JBJS Am 1983;65(9):1232–1244, PMID 6654936
26
Assessment of Spine Surgery Outcomes: Inconsistency of Change Amongst Outcome Measurements
Spine J · 2010 · Copay AG, Martin MM, Subach BR, Carreon LY, Glassman SD, Schuler TC, Berven S · Multicentre PROM analysis, n=460
What it does: Applies four validated outcome measures to the same 460 lumbar surgery patients, each judged against its own established minimal clinically important difference. Only 40.5% of patients were classified consistently as improved or not improved across all four.
Epistemological angle: "Clinically meaningful improvement" turns out to be a property of the instrument rather than of the patient. The same person, at the same visit, is a success on one validated measure and a failure on another. Which means the choice of primary outcome in any trial is a covert decision about who will count as helped — a decision made before any data are collected, usually without discussion, and capable of determining the result.
My take: Read this before the next trial you assess on the strength of its primary endpoint. The endpoint was chosen by people who knew roughly what their intervention was likely to move.
27
Surgical vs Nonsurgical Treatment of Adults with Displaced Proximal Humeral Fractures (PROFHER)
JAMA · 2015 · Rangan A, Handoll H, Brealey S et al. · Pragmatic multicentre RCT, n=250
What it does: Compares surgical fixation or replacement against sling immobilisation for displaced proximal humeral fractures involving the surgical neck. No significant difference in Oxford Shoulder Score or quality of life at two years, and similar complication rates. The authors state that the results do not support the trend toward increased surgery.
Epistemological angle: Included here rather than among the trials because of what preceded it. Surgical volume for these fractures had been climbing for a decade on the strength of anatomical reasoning — restore the anatomy and function follows — without a trial supporting it. Mechanistic plausibility functioned as sufficient warrant to adopt, and only a large pragmatic trial was accepted as sufficient warrant to reconsider. The two standards are not the same size.
My take: Restore the anatomy, restore the function is one of the most reasonable-sounding sentences in orthopaedics and it is a hypothesis, not a principle. It has been tested a number of times now and its record is mixed at best.
Also worth reading
Fleming TR, DeMets DL. Surrogate end points in clinical trials: are we being misled? Ann Intern Med 1996;125(7):605–613 · PMID 8815760 · doi:10.7326/0003-4819-125-7-199610010-00011. States the full-mediation criterion formally, and assembles the cases entry 23 dramatises. Most surrogates in orthopaedics have never been assessed against this standard and a good many would fail it.
Theme Six
What observation can and cannot buy
What does allocation buy you, and what can a cohort do without it?
Randomisation is often treated as a property a trial either has or lacks. It is better understood as a loan against future adherence: when patients and surgeons cross between arms in large numbers, the loan is called in. But the converse lesson is the one this theme exists to correct. Observational evidence is not a weaker version of a trial — it answers questions no trial can reach, and where it fails it usually fails for a specific and correctable reason rather than because it lacked randomisation.
28
SPORT: The Trial and Its Observational Twin
JAMA · 2006 · Weinstein JN, Tosteson TD, Lurie JD et al. · RCT (n=501) and parallel observational cohort (n=743), published the same day
What it does: Randomises patients with lumbar disc herniation to surgery or non-operative care. By three months only half the surgical arm had been operated on and 30% of the non-operative arm had crossed to surgery. The intention-to-treat analysis showed small, non-significant differences, and the authors state that conclusions about superiority or equivalence are not warranted from it. The companion paper follows those who declined randomisation, analysed as treated, and finds a clear significant advantage for surgery on every primary outcome.
Epistemological angle: The best teaching pair in this collection. Same investigators, same outcome measures, same eligibility criteria, same journal, same day — and abandoning randomisation for observed treatment choice produces a dramatically cleaner and more confident-looking answer. Which of the two papers a surgeon cites is a reliable indicator of what they believed before opening the journal. The observational result is not fabricated; it is what confounding by indication looks like when the people choosing surgery are systematically different from those who do not.
My take: A well-funded, well-run, adequately powered trial that cannot answer its own question. That is worth sitting with, because the usual explanations for disappointing trials — too small, too sloppy, wrong outcome — do not apply here. Adherence is a precondition for randomisation to mean anything, and in surgical trials it is frequently unobtainable.
29
Hormone Therapy and Coronary Heart Disease: The Cohort, the Trial, and the Reconciliation
JAMA · 2002 · Rossouw JE, Anderson GL, Prentice RL et al. (Women's Health Initiative), RCT n=16,608 · with Epidemiology · 2008 · Hernán MA, Alonso A, Logan R, Grodstein F, Michels KB, Willett WC, Manson JE, Robins JM · re-analysis of the Nurses' Health Study
What it does: A large, careful, decades-long observational cohort had found that women taking combined hormone therapy had less coronary heart disease. The randomised trial found the opposite: a hazard ratio of 1.29 for coronary events, with stroke and pulmonary embolism also raised, and the trial stopped early because harms exceeded benefits. Six years later Hernán and colleagues re-analysed the original cohort as though it were a sequence of trials, emulating the randomised design and its intention-to-treat analysis, and largely reproduced the trial's estimates from the same observational data.
Epistemological angle: Included as the general-medicine counterpart to entry
28, and it teaches something that pair cannot. The first two papers look like the familiar story that observational evidence is unreliable and randomisation is the remedy. The third undoes that reading. The cohort's data were not wrong; the analysis had compared prevalent users against never-users, which selects for women who had already tolerated the drug, and had aligned time zero incorrectly. Re-analysed with the trial's own design imposed on it, the cohort gave the trial's answer. So the failure was not observational data as a category but a specific and correctable analytic choice — which means the lesson is not "distrust cohorts" but "ask what trial this analysis is emulating, and whether anyone specified it." That is a far more useful question to carry into a registry paper, and it is the question entry
30 passes and most do not.
My take: This is the entry I would put in front of anyone who says the words "real-world evidence" approvingly. The gap between the cohort and the trial was not a story about randomisation's magic. It was a story about a design question nobody had been required to ask out loud, and about how long it took — six years and a good methodologist — to find out that the data had been able to answer correctly all along.
30
Failure Rates of Stemmed Metal-on-Metal Hip Replacements: Analysis of the National Joint Registry of England and Wales
Lancet · 2012 · Smith AJ, Dieppe P, Vernon K, Porter M, Blom AW · Registry cohort, n=402,051 primary hip replacements
What it does: Analyses four hundred thousand hip replacements and finds that stemmed metal-on-metal implants failed at markedly higher rates than alternatives, with risk rising with head size and concentrated in younger women.
Epistemological angle: The positive case for observational data, and it turns entirely on class size. The signal was slow, dose-dependent and modified by patient characteristics — invisible to any individual surgeon's experience and beyond the power of any feasible randomised trial. The effect modification by sex and head size is the part a trial would most likely have missed even if one had existed. Corroboration from the Australian registry on a different population is what converted a worrying pattern into a conclusion. Note the precondition the next two entries remove: the registry is compulsory, national, and held by nobody who sells implants.
My take: Worth holding alongside the sham trials as a corrective. The lesson of this collection is not that observational evidence is weak and trials are strong. It is that each answers questions the other cannot, and knowing which question you are asking is the whole skill.
Lancet 2012;379(9822):1199–1204 · PMID 22417410 ·
doi:10.1016/S0140-6736(12)60353-5 · corroborated by de Steiger RN et al., JBJS Am 2011;93(24):2287–2293, PMID 22258775
Also worth reading
Montori VM, Guyatt GH. Intention-to-treat principle. CMAJ 2001;165(10):1339–1341 · PMID 11760981 · PMC81628 · no DOI in record. Short, free, and the reason entry 28's two papers are one valid comparison and one that has quietly reverted to observational status. Intention-to-treat is not a conservative convention but the condition under which a randomised comparison retains any causal warrant at all.
Katz JN, Brophy RH, Chaisson CE et al. Surgery versus physical therapy for a meniscal tear and osteoarthritis (METEOR). N Engl J Med 2013;368(18):1675–1684 · PMID 23506518 · doi:10.1056/NEJMoa1301408. Crossover designed in rather than suffered: 30% of the physiotherapy arm had surgery within six months, so the null is equally consistent with equivalence and with a third of that arm having received the comparator.
Theme Seven
What survives the retelling
What survives the compression from data to conclusion sentence, and who decided what went in?
Almost nobody reads trials. Surgeons read abstracts, and mostly the last sentence of them. This theme follows one continuum from end to end: outcomes forgotten between protocol and paper, conclusions framed to matter, hedges present in the Discussion and absent from the Abstract, and — at the far end, where the same failure stops looking accidental — data the reader never had the opportunity to see at all. The intensity varies. The structure does not.
31
Empirical Evidence for Selective Reporting of Outcomes in Randomized Trials
JAMA · 2004 · Chan AW, Hróbjartsson A, Haahr MT, Gøtzsche PC, Altman DG · Protocol-to-publication cohort, 102 trials, 3736 outcomes
What it does: Compares ethics-committee-approved protocols against the papers eventually published. Half of efficacy outcomes and 65% of harm outcomes were incompletely reported. Statistically significant outcomes were two to five times more likely to be fully reported. 62% of trials had at least one primary outcome changed, added or dropped. When the trialists were surveyed, 86% denied that unreported outcomes existed.
Epistemological angle: The field-defining paper, and the last figure is the one that matters. The investigators were not concealing results; they had forgotten. Selective reporting is mostly not misconduct but ordinary cognition operating on data the analyst has already seen — outcomes that behaved interestingly feel like the real findings, and the others fade. This is why procedural remedies like pre-registration are necessary: the failure mode is invisible from inside the person committing it.
My take: It seems safest to assume the published literature leans towards benefit and away from harm — not through anyone's dishonesty, but because that is the shape of the filter it passes through.
32
Analyzing Spin in Abstracts of Orthopaedic RCTs with Statistically Insignificant Primary Endpoints
Arthroscopy · 2020 · Arthur W, Zaaza Z, Checketts JX et al. · Meta-research, 250 RCTs screened
What it does: Examines abstracts of orthopaedic trials whose primary endpoint was null. 44.8% contained spin; among those, 79.5% had it in the conclusion. JBJS showed the highest prevalence at 56.8%. There was no association with industry funding.
Epistemological angle: Two findings, and the second is the interesting one. Spin concentrates in the conclusion sentence — the part actually read, quoted and remembered — so the distortion is maximal at the point of highest readership. And the absence of a funding association undercuts the comfortable story that this is a money problem. It is a wanting-your-work-to-matter problem, which is far more widespread and much harder to legislate against.
My take: Note that this sits awkwardly beside a wider meta-research literature in which industry funding does predict favourable conclusions. I have not resolved the tension and would rather leave it standing than pick a side: competent groups asking adjacent questions have got different answers about the role of money.
33
Epistemic Asymmetry Between Abstracts and Discussion Sections in the Orthopaedic Literature
JBJS · 2026 · Parisien R · Quantitative linguistic corpus analysis, 201 JBJS and CORR publications
What it does: Measures the epistemic stance of the same paper in two places. Across 201 orthopaedic publications, Abstracts hedge markedly less than the Discussion sections of the papers they summarise, with a large effect size (Cohen's d = 1.15). A companion measure found that 94% of Abstracts surface no limitation at all. Open-access and paywalled papers did not differ, which locates the asymmetry in the genre rather than in who can read it.
Epistemological angle: Entry
32 shows that some abstracts overstate null results. This asks a prior question of every abstract, null or positive: how much of the paper's own stated uncertainty survives into the part that gets read. The answer is that the qualifications are written and then removed at the point of compression — the authors know what their study cannot support and say so where almost nobody looks. That reframes the problem. It is not that the field fails to acknowledge uncertainty, which entry
36 might be read as showing; it is that acknowledgement is structurally quarantined in the section with the smallest readership. Whether that constitutes a distortion depends on what one thinks an Abstract is for, and the paper does not settle that question.
My take: The 94% is the number worth arguing about, and the obvious objection is that an Abstract has a word limit and is not the place for caveats. That is fair as far as it goes, but the same limit does not prevent a stated effect size, so the omission is a choice about what the space is for. The study measures hedging language rather than whether readers were actually misled, which is the question one would want answered and a harder one to design. What I can say is that writing it changed how I read my own abstracts.
J Bone Joint Surg Am 2026 · in press; volume, pages, PMID and DOI to be supplied at publication · a commissioned Commentary & Perspective accompanies it
34
VIGOR and the Expression of Concern
NEJM · 2000 · Bombardier C, Laine L, Reicin A et al. (VIGOR), RCT n=8,076 · and NEJM · 2005 · Curfman GD, Morrissey S, Drazen JM, Expression of Concern
What it does: VIGOR reported that rofecoxib caused fewer serious gastrointestinal events than naproxen, and noted in passing a higher rate of myocardial infarction in the rofecoxib arm (0.4% against 0.1%), attributed in discussion to a protective effect of naproxen. Five years later the journal's editors published an Expression of Concern stating that additional myocardial infarctions known to at least some authors before publication had not appeared in the paper.
Epistemological angle: The custody problem stated outside orthopaedics, and placed before the two entries that state it inside. The failure here is not that peer review missed a subtle signal; the harm data existed and the reviewing journal did not have them. That locates the vulnerability precisely: peer review evaluates a manuscript, and a manuscript is a selection from a dataset made by a party with an interest in the selection. Note also the asymmetry of framing within the paper itself — the same numerical contrast was read as naproxen protecting rather than rofecoxib harming, which is a live demonstration of how much work the choice of comparator does when a result is turned into a sentence.
My take: I include this partly so that entry
35 is not read as a story about orthopaedics. It is a story about what happens when the party that owns the data also owns the product, and that arrangement is not ours. Ours is simply one of the places it has been documented.
35
The BMP-2 Sequence: What the Published Record Showed and What the Data Showed
Spine J · 2011 · Carragee EJ, Hurwitz EL, Weiner BK, comparison of 13 industry-sponsored trials against regulatory data · and Ann Intern Med · 2013 · Fu R, Selph S, McDonagh M et al., independent individual-patient-data reanalysis
What it does: Carragee and colleagues compared what thirteen industry-sponsored trials of rhBMP-2 reported against what regulatory submissions and later follow-up showed. None of the thirteen published papers reported any adverse event attributable to the product; comparison with other sources suggested true rates ten to fifty times higher. Two years later Fu and colleagues obtained the underlying individual patient data through the Yale Open Data Access project and reanalysed it independently, finding no clinical advantage over iliac crest bone graft and increased harms including cancer risk, with the early publications having misrepresented both.
Epistemological angle: The most disturbing pair here and a necessary one. Peer review functioned, publication proceeded, specialty consensus formed, and the harm signal sat in a regulatory file nobody was reading. "Published in a good journal" is a statement about a process, not a warrant of truth, and the process has a specific blind spot: it evaluates the manuscript in front of it rather than the data behind the manuscript. What eventually settled the dispute was not more argument about the published papers, more editorials, or more expert opinion, all of which had been tried. It was independent access to the raw data. Transparency in this story is not an ethical nicety appended to good science; it was the only mechanism capable of terminating the disagreement.
My take: Thirteen trials and zero adverse events. With hindsight that figure is hard to credit, since no surgical intervention is without complications; at the time it read as a good safety profile, which is how such figures usually read. The implication of the pair is uncomfortable and I have not found a way to soften it: where the sponsor controls the data, the published literature may not be sufficient evidence on its own, however many trials it contains.
Also worth reading
Okike K, Kocher MS, Wei EX, Mehlman CT, Bhandari M. Accuracy of conflict-of-interest disclosures reported by physicians. N Engl J Med 2009;361(15):1466–1474 · PMID 19812403 · doi:10.1056/NEJMsa0807160. 71.2% of payments disclosed overall but only 50.0% of indirectly related ones. A disclosure statement reports which relationships the author judged relevant, which is a different and less useful thing than which existed.
Ruelos VCB, Masood R, Puzzitiello RN et al. The reverse fragility index. Knee Surg Sports Traumatol Arthrosc 2023;31(8):3412–3419 · PMID 37093236 · doi:10.1007/s00167-023-07420-0. Median three outcome events would flip a reported null to significant, and in 81.3% of studies loss to follow-up exceeded that number.
Theme Eight
The surgeon as instrument
What happens to evidence when the treatment is a human being?
A drug is the same molecule in every trial. A surgical procedure is a person, on a particular day, at a particular point on their learning curve, classifying the pathology by eye. Everything orthopaedics borrowed from pharmaceutical trial methodology inherited an assumption that does not hold here. This theme is about the instrument: how well it is calibrated, how it perceives, the trial designs built around its variability, and what it turns out the operation was doing when it worked.
36
Do Orthopaedic Surgeons Acknowledge Uncertainty?
CORR · 2016 · Teunis T, Janssen S, Guitton TG, Ring D, Parisien R · Cross-sectional survey, n=242
What it does: Measures surgeons' recognition of uncertainty and looks for what predicts it. Recognition was uniformly low and did not improve with years in practice, while confidence bias increased with experience. Better statistical understanding was the strongest independent predictor of acknowledging uncertainty. Greater trust in the orthopaedic evidence base was independently associated with acknowledging less.
Epistemological angle: That last association is the finding. Surgeons who believe orthopaedics rests on solid evidence are less able to say "I don't know" — which is exactly backwards, given that the honest reading of the preceding thirty-four entries is that the evidence base does not support such confidence. The paper documents a specialty whose self-assessment runs contrary to its actual epistemic position. That is not a knowledge deficit; it is a calibration failure, and calibration failures are by their nature invisible from the inside. Read with entry
33, which suggests the acknowledgement may exist and be getting stripped out at the point of publication rather than never forming.
My take: Three criticisms, and they strengthen it. The instrument measures self-reported acknowledgement rather than actual calibration — a surgeon may score well and still judge badly. The design is cross-sectional, so rising confidence with experience is equally explicable as a cohort effect. And the sample came from the Science of Variation Group, meaning it drew on the most uncertainty-aware population in the specialty, and recognition was still uniformly low. Every limitation points the same way as the conclusion. I wrote this one, and I think the real number is worse than we reported.
Clin Orthop Relat Res 2016;474(6):1360–1369 · PMID 26552806 ·
doi:10.1007/s11999-015-4623-0 · erratum 474(6):1530–1531, PMID 26861152; Editor's Spotlight, Leopold SS, 474(6):1356–1359, PMID 26818597
37
Two Proposals for Trials Whose Instrument Is a Person: IDEAL and Expertise-Based Randomisation
Lancet · 2009 · McCulloch P, Altman DG, Campbell WB et al., consensus framework · and BMJ · 2005 · Devereaux PJ, Bhandari M, Clarke M, Montori VM, Cook DJ, Yusuf S, Sackett DL et al., methodological proposal
What it does: IDEAL proposes a five-stage framework for evaluating surgical innovation — idea, development, exploration, assessment, long-term study — built explicitly around confounding by operator, team, learning curve and perceived equipoise. Devereaux and colleagues propose a change to allocation itself: randomise patients to a surgeon expert in procedure A or a surgeon expert in procedure B, rather than to a procedure that whichever surgeon is available then performs.
Epistemological angle: Two responses to the same structural problem, and they are complementary rather than alternative. IDEAL states the problem: a surgical treatment is not a stable standardisable input the way a drug is, so a trial design imported wholesale from pharmacology answers a question about an intervention that does not exist in the form assumed. The learning curve alone means a procedure evaluated early is a different intervention from the same procedure evaluated late, and conventional analysis treats them as one. Expertise-based randomisation is the design consequence: if the same surgeon delivers both arms and is better at one of them, the trial measures A against B as performed by people trained mostly in A. Making the surgeon part of the allocated intervention is the right response to a variable that is not noise but structure — at the cost, which the authors accept, that the result then attaches to a procedure-plus-practitioner package rather than to a procedure.
My take: Both are widely cited and neither is much followed. Orthopaedic innovation still arrives fully formed in a case series, which says something about incentives rather than about the frameworks. The expertise-based design has an obstacle worth naming honestly, because it is not logistical: it requires a surgeon to concede in advance and in writing that a colleague is better at something. That is a real cost, and pretending the barrier is purely practical has not helped anyone adopt it in twenty years.
38
The Invisible Gorilla Strikes Again: Sustained Inattentional Blindness in Expert Observers
Psychol Sci · 2013 · Drew T, Võ ML-H, Wolfe JM · Experimental study, 24 radiologists
What it does: Asks radiologists to perform a familiar lung-nodule detection task on chest CT. An image of a gorilla, forty-eight times the size of the average nodule, was inserted into the final case. Eighty-three per cent of the radiologists did not report it. Eye tracking showed that most of those who missed it had looked directly at it.
Epistemological angle: The front matter of this collection asserts that observation is theory-laden. This is the entry that demonstrates it, and it demonstrates something stronger than the usual formulation. The claim is not merely that expectation colours interpretation of what was seen; it is that expertise organises the visual search itself, so that the framework determines what enters awareness at all. The gorilla was fixated and not seen. That places the effect upstream of judgment, where no amount of care in reasoning can reach it, and it is why entry
49's finding about classification disagreement should not be read as carelessness. A skilled observer is a tuned instrument, and tuning is subtraction as well as amplification.
My take: This is the one I like to point out, because it is the only entry here that people find funny, and the laugh is the point at which the argument lands. The thing worth taking away is that expert perception is constructed, not that any particular specialty is careless. It also cuts against a comfort I notice in myself: I read a scan differently having formed an impression from the history, and I have always filed that under experience rather than under the thing this paper measures.
39
Outcomes After Rotator Cuff Repair Are Not Related to Structural Healing
Arthroscopy · 2021 · Holtedahl R, Bøe B, Brox JI · Meta-regression, 64 RCT and 19 cohort arms
What it does: Pools trial arms of rotator cuff repair. Retear rate was around 20% at a median of thirteen months. Age and tear size predicted retear. Clinical outcome did not track whether the repair had actually healed.
Epistemological angle: This has something close to the shape of a Gettier case, and it is worth being careful about how close. The belief "rotator cuff repair helps patients" is true. The justification every surgeon holds — the torn tendon is reattached, it heals, function follows — is disconnected from the truth-maker in a substantial fraction of the people who benefit. Whether that is Gettier proper or simply a true belief resting on a false explanation is arguable, and the arguable part is instructive rather than a defect. Either way the practical upshot holds: the reason the belief is true is not the reason we have for holding it. The same structure applies to Moseley and vertebroplasty, where patients genuinely improved and the stated mechanism was doing none of the work.
My take: If a fifth of repairs fail structurally and those patients do as well, then something other than the repair is producing the benefit and we do not know what it is. That is a more interesting research question than another comparison of fixation constructs.
Theme Nine
When the instrument is a machine
What kind of witness is a machine, and on which population was it right?
The organising claim of this theme is that artificial intelligence raises almost no epistemological problem the preceding entries have not already raised. Shortcut learning is a reference class failure. Automation bias is a calibration failure. The argument about explainability is the argument about mechanism and outcome. What is new is that the errors are distributed differently from ours and arrive in a confident register, which makes them harder for us to catch than our own. The theme opens with the best evidence any of this has, because a reading list assembled out of failures would flatter our scepticism and teach nothing.
40
Mammography Screening with Artificial Intelligence (MASAI)
Lancet Oncol 2023, Lancet Digit Health 2025, Lancet 2026 · Lång K, Hernström V, Gommers J et al. · Randomised, controlled, population-based screening-accuracy trial, n=105,934
What it does: Randomises women in the Swedish national screening programme to AI-supported screen reading or standard double reading. The AI triaged examinations to single or double reading and marked suspicious findings; radiologists still read every case. Cancer detection was 6.4 against 5.0 per 1000 (ratio 1.29) with no increase in false positives. The primary endpoint, reported in 2026, was interval cancer rate: non-inferior at 0.88, with sensitivity 80.5% against 73.8% (p=0.031) and specificity identical at 98.5%. Screen-reading workload fell by 44%.
Epistemological angle: This is the entry that holds AI to the standard the rest of this collection demands of everyone else, and it is worth being explicit that it meets that standard where most of the sham-tested procedures in theme one did not. Randomised, population-based, pre-registered, primary endpoint pre-specified and reported on a hard outcome rather than a surrogate. Note what was actually tested: not the machine replacing the radiologist but the two together against the standard of care. The comparator matters as much as the result — Swedish double reading is a demanding benchmark, and the intervention beat it on sensitivity while halving the human reading burden. What licenses the result is also worth naming: the model's training distribution and its deployment population were the same, which is precisely the condition entry
42 shows to be absent when these systems fail.
My take: I include this first in the theme deliberately. It would be easy to assemble an AI reading list out of failures, and the result would be a collection that flatters our scepticism and teaches nothing. The honest position is that the best available evidence for AI-supported reading is stronger than the evidence for a good deal of what I do in theatre, and I would rather say so plainly than bury it.
41
An Algorithmic Approach to Reducing Unexplained Pain Disparities in Underserved Populations
Nat Med · 2021 · Pierson E, Cutler DM, Leskovec J, Mullainathan S, Obermeyer Z · Deep learning on knee radiographs
What it does: Trains a model on knee radiographs to predict the patient's reported pain rather than to reproduce a radiologist's severity grade. Standard radiographic grading accounted for 9% of unexplained disparity in pain between served and underserved patients; the algorithmic prediction accounted for 43%. The authors conclude that much of the unexplained pain arises from features within the knee that standard radiographic measures do not capture, and note that since severity measures drive treatment decisions, this bears directly on access to arthroplasty.
Epistemological angle: An orthopaedic classification in use since 1957 was shown to be lossy by a system with no stake in it. That is worth separating from the usual framing. The finding is not that the algorithm is cleverer; it is that the radiographic grade discards information the knee contains, and that what it discards is not randomly distributed. Read against entry
47, this is the same lesson from the opposite direction: an ordinal category imposed on a continuous structure loses signal, and the loss is invisible from inside the category because the category is what you are looking with. It is also the collection's clearest case of a machine functioning as a genuinely third kind of cognizer — not faster at our task, but sensitive to something our task was never set up to see.
My take: This is the paper that changed my mind about what these systems are for. I had thought of them as accelerating what we already do. Here the value came from asking a different question — predict the symptom, not the grade — and the answer implied that our grading scheme was the problem. I would like to see this design repeated in the shoulder and the spine, and I do not know why it has not been.
42
Deep Learning Predicts Hip Fracture Using Confounding Patient and Healthcare Variables
npj Digit Med · 2019 · Badgeley MA, Zech JR, Oakden-Rayner L et al. · 17,587 pelvic radiographs
What it does: Trains models to classify hip fracture and, separately, five patient traits and fourteen hospital process variables. All twenty could be predicted from the radiograph, with the best performance on scanner model (AUC 1.00), scanner brand (0.98) and whether the order was marked priority (0.79). Fracture prediction reached AUC 0.78 from the image alone and 0.91 with patient and process data added. On a test set balanced across patient and process variables the model performed at chance, AUC 0.52.
Epistemological angle: The reference class problem of entry
17 in its purest mechanical form, and the reason that entry belongs upstream of this whole theme. A trained model's reference class is its training distribution, and nothing in the model selects that class — the data do, silently, through whatever happens to covary with the label. Here the class turned out to be something like "pelvic radiographs from this scanner, at this hospital, ordered with this priority flag." The 0.52 is the number to carry: strip the confounders and the fracture detector detects nothing. Hájek's point that the choice cannot be eliminated but only made visible applies exactly, and it is why the deployment question for these systems is not how accurate they are but on which population that accuracy was estimated, and whether the patient in front of you is in it.
My take: Set beside entry
40, which works, this stops being a cautionary tale and becomes the criterion. MASAI's population and its training population were the same. This model's were not, and nobody knew because nobody had asked what class it had learned. That is a question a surgeon can ask of a vendor without any machine learning at all, and it is the only one I would insist on.
43
Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance
Radiology · 2023 · Dratsch T, Chen X, Rezazade Mehrizi M et al. · Prospective experiment, 27 radiologists, 50 mammograms
What it does: Radiologists of varying experience read mammograms with the aid of a purported AI system that supplied a deliberately incorrect BI-RADS category for twelve of forty cases. Correct assessments fell from 79.7% to 19.8% among inexperienced readers, from 81.3% to 24.8% among the moderately experienced, and from 82.3% to 45.5% among the very experienced.
Epistemological angle: The counterpart to entry
40, and the reason a favourable trial does not settle the question. MASAI shows what a well-matched system adds to a reader; this shows what a wrong one subtracts, and the subtraction is larger than the addition. The epistemological content is about testimony. A machine's output enters the reader's judgment not as one datum among several but as an anchor, and the very experienced group — down thirty-seven points — demonstrates that expertise is protection rather than immunity. Read with entry
19: the failure is in the same family as the anchoring effects that moved surgeons' answers by 22 points, except that here the anchor arrives with institutional authority attached.
My take: The design is artificial and the effect size is therefore an upper bound; nobody deploys a system that is wrong 30% of the time. What survives the artificiality is the direction and the gradient across experience, and those are enough. If I am going to use these tools, I should want to know my own susceptibility, and I have no way to measure it.
44
Endoscopist Deskilling Risk After Exposure to Artificial Intelligence in Colonoscopy
Lancet Gastroenterol Hepatol · 2025 · Budzyń K, Romańczyk M, Kitala D et al. · Multicentre observational study, 1,443 patients, four Polish centres
What it does: Compares the adenoma detection rate of standard, unassisted colonoscopy in the three months before AI detection tools were introduced against the three months after. Detection on unassisted procedures fell from 28.4% to 22.4%, an absolute difference of 6.0% (adjusted odds ratio 0.69). The comparison is of the endoscopists' own unaided performance, before and after they became habituated to assistance.
Epistemological angle: This is the theme's most important entry for anyone thinking about how to integrate these systems, and it is not an argument against them. Entry
40 establishes that the combination outperforms the human alone. This establishes that the human inside the combination may not be the human who entered it. Both can be true, and taken together they reframe the question from whether to adopt to what to measure afterwards — the unaided performance of the assisted clinician is a quantity nobody was tracking, and it turns out to move. The design cannot establish causation and the authors do not claim it does; it is a before-and-after comparison with all the vulnerabilities that implies.
My take: Six percentage points of adenoma detection is not a small number, and it was found by people who were running an AI trial and thought to look sideways. That is the part I find instructive. The effect was available to be measured for as long as these tools have been deployed, and it was measured once, late, and by accident of study design. Whatever the true magnitude, the more durable finding is how nearly we did not look.
45
Ghassemi versus London: Must a Clinical Machine Be Able to Explain Itself?
Lancet Digit Health · 2021 · Ghassemi M, Oakden-Rayner L, Beam AL · and Hastings Cent Rep · 2019 · London AJ · Opposing arguments
What it does: Ghassemi and colleagues argue that current explainability methods do not deliver what is claimed for them — they produce post-hoc rationalisations that can fail silently for the individual patient — and that rigorous internal and external validation is the more direct route to the goals explainability is supposed to serve. London argues that opaque decisions are far more common in medicine than critics acknowledge, and that where causal understanding is thin, the ability to produce and empirically verify results can matter more than the ability to explain how they were produced.
Epistemological angle: Taught as a pair, in the manner of a genuine disagreement between competent people rather than a controversy with a right answer. The dispute is the mechanism-versus-outcome argument this collection has already run twice, transposed. Entry
23 shows a drug that suppressed the mechanism and killed patients; entry
39 shows repairs that fail structurally in patients who do well. In both, the causal story we could tell was disconnected from the result we could verify. London's argument is that this is medicine's normal condition and that demanding explanation from machines while tolerating its absence in ourselves is inconsistent. Ghassemi's is that the inconsistency does not license accepting an explanation we know to be confabulated. Both are correct about something, and the collection cannot resolve it.
My take: London's argument is the uncomfortable one for me, because it is continuous with my own material. I have spent a good deal of this collection arguing that we reason from mechanism when we should reason from outcome, and it would be inconsistent to reverse that the moment the reasoner is a machine. Where I part from him is on what verification requires: entry
42 shows that a model can be verified impeccably on the wrong population, and no amount of outcome data from that population tells you so.
46
Can Artificial Intelligence Pass the American Board of Orthopaedic Surgery Examination?
CORR · 2023 · Lum ZC · Benchmark study, 207 in-training examination questions
What it does: Puts orthopaedic in-training examination questions to a large language model and benchmarks against five years of resident performance data. 47% correct, roughly first-year level, below the tenth percentile of final-year residents — with accuracy falling sharply as questions moved from recall toward application.
Epistemological angle: The right frame is testimony, not competence. The question is not whether the machine is clever but under what conditions its output may be taken on trust, and the answer here is precise: it degrades exactly where clinical judgment lives. A source that is reliable on recall and unreliable on synthesis is a genuinely difficult kind of witness, because its failures arrive in the same confident register as its successes. The specific numbers will date within months. The framing will not.
My take: Much of this collection concerns the ways our calibration, our use of base rates, and our consistency fall short of the ideal. It would be convenient to conclude that machines should therefore take over, and this paper is a caution against that: the machine's errors are differently distributed, not absent, and we are worse at detecting them because they do not look like human mistakes.
Also worth reading
Zech JR, Badgeley MA, Liu M, Costa AB, Titano JJ, Oermann EK. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs. PLoS Med 2018;15(11):e1002683 · PMID 30399157 · doi:10.1371/journal.pmed.1002683. The general case behind entry 42: sorting by hospital alone achieved AUC 0.861 because pneumonia prevalence differed between sites by a factor of thirty, and the networks identified the source institution in 99.95% of images.
Theme Ten
Forced categories
What happens when a continuous or subjective phenomenon is made to be one thing or another?
This theme is last because it is first. Every preceding theme presupposes that there is a kind here to have base rates about, to randomise treatments for, to measure with surrogates and to argue over. Where the category does not carve — where union, impingement, degeneration or clinically meaningful improvement is a boundary drawn across a continuum rather than a joint the world already has — nothing downstream repairs it. You cannot compute a prevalence for a thing that is not there to be prevalent, and no larger trial will settle whether it is. These are questions of category, vagueness and identity, and they are not empirical. It is also the least developed area in this collection, which is why the marginal return on work here seems to me higher than anywhere else on the list.
47
The Cost of Dichotomising Continuous Variables
BMJ · 2006 · Altman DG, Royston P · Statistical note
What it does: One page, setting out what is lost when a continuous measurement is converted into two categories. Information is discarded; a step change in risk is assumed where the data show a gradient; individuals close to the cutpoint on opposite sides are treated as different in kind while those at opposite ends of the same category are treated as the same; statistical power falls; and where the cutpoint is chosen after inspecting the data, the resulting effect estimate is optimistic.
Epistemological angle: The statistical statement of this theme's thesis, and deliberately placed first because it is the cheapest to grasp and the hardest to unsee. Once the argument is in view it applies to almost every categorical variable in orthopaedic research. Union and non-union is a dichotomy imposed on a continuous process of mineralisation. Success and failure on a patient-reported outcome is a dichotomy imposed on a continuous distribution of experience. Displaced and undisplaced is a dichotomy imposed on a continuum of millimetres. In each case the categorical variable is the thing we count, analyse and report, and Altman and Royston's point is that the choice of where to cut is doing work that the data are then held to have done.
My take: The last of their objections is the one I would emphasise for our field. A cutpoint chosen after the results are known is not a measurement decision but an inference dressed as one, and radiographic union thresholds have largely been arrived at that way. This is a one-page paper and I have never seen it cited in an orthopaedic manuscript.
48
Guidance for Modifying the Definition of Diseases: A Checklist
JAMA Intern Med · 2017 · Doust J, Vandvik PO, Qaseem A et al. · Consensus checklist, thirteen-member multidisciplinary group
What it does: Observes that guideline panels frequently widen disease definitions, increasing the proportion of a population labelled unwell, and that no guidance existed for panels considering such a change. Proposes an eight-item checklist covering the number of people affected by the change, the trigger for it, the prognostic ability and precision of the revised definition, and the balance of benefits and harms — with the explicit requirement that a definitional change be justified rather than announced.
Epistemological angle: The normative counterpart to entry
47. Altman and Royston show what dichotomising costs; this treats the placement of the cut as a decision requiring argument, which is precisely the move the orthopaedic literature has not made. The orthopaedic instance is the osteoporosis T-score: a threshold at 2.5 standard deviations below a young-adult reference mean, applied to a continuous and unimodal distribution of bone density, which brought a disease into existence by stipulation and set the denominator for everything subsequently claimed about treating it. Note also what the checklist implies about entries
8 to
10 — if a definitional change alters how many people are called unwell, then the label's effects documented there are a consequence of a decision somebody made, not a fact about the world.
My take: I would like to see a version of this checklist applied retrospectively to the definitions we use every day, beginning with union. My expectation is that most of them would fail item four — precision and accuracy of the definition — not because they are poorly made but because nobody was ever required to state what the definition was for.
49
The Neer Classification System for Proximal Humeral Fractures: Interobserver Reliability and Intraobserver Reproducibility
JBJS Am · 1993 · Sidor ML, Zuckerman JD, Lyon T, Koval K, Cuomo F, Schoenberg N · Observer-agreement study, 50 fractures, 5 observers
What it does: Has five observers of varying seniority classify the same fifty fractures twice, six months apart. Interobserver kappa 0.48 to 0.52. All five agreed on roughly 30% of fractures. Collapsing the system from sixteen categories to six did not improve agreement.
Epistemological angle: A classification can look like objective anatomical description while functioning as a low-reliability perceptual judgment. Every study that stratifies treatment or outcome by Neer type silently inherits this noise, which means a substantial literature is built on a variable that five experts agree about less than a third of the time. That simplification did not help is the diagnostic detail: the problem is in the seeing, not the taxonomy.
My take: A parallel from the hip is worth putting beside this. Frandsen's 1988 work on the Garden classification found agreement in 22 of 100 fractures, and the field responded by collapsing Garden into displaced versus undisplaced. Coarsening did not rescue Neer, so it is not a general remedy — but where the coarser distinction is the one that actually drives treatment, it is honest about how much the observer can see. That is one of the few clean self-corrections in this collection.
50
Development of the Radiographic Union Score for Tibial Fractures (RUST)
J Trauma · 2010 · Whelan DB, Bhandari M, Stephen D, Kreder H, McKee MD, Zdero R, Schemitsch EH · Reliability study, 45 radiograph sets, 7 reviewers
What it does: Establishes that a radiographic scoring system for tibial fracture healing is reproducible — interobserver ICC 0.86, intraobserver 0.88. The authors state plainly that no gold standard for union exists against which the score has been validated.
Epistemological angle: The cleanest available separation of reliability from validity. Seven observers agreeing tells you the instrument is consistent; it tells you nothing about whether it measures a healed, painless, usable limb. RUST has since become the most widely adopted union scale in trauma research and functions as a de facto trial endpoint. The caveat the authors printed has not travelled with it.
My take: This is not a criticism of the paper, which did exactly what it claimed and flagged what it had not done. It is a criticism of what a field does with a well-made instrument: reliability is easy to demonstrate and validity is hard, so we demonstrate reliability and then quietly start behaving as though we had shown the other thing.
51
Rotator Cuff Related Shoulder Pain: Assessment, Management and Uncertainties
Man Ther · 2016 · Lewis J · Position paper
What it does: Argues that "subacromial impingement syndrome," "rotator cuff tendinopathy" and the tear categories should be collapsed into a single umbrella term that claims less — rotator cuff related shoulder pain — on the grounds that the existing labels assert a causal mechanism that has not been shown and that imaging findings correlate poorly with symptoms.
Epistemological angle: The only entry here that engages the nosological question directly: not "does this treatment work" but "does this named entity correspond to anything real enough to be a diagnosis." Lewis had made the same argument eight years earlier and the field had not moved; taken together the two papers show a specialty revising a category in real time, which is a slow and visible process rarely captured in the literature.
My take: Renaming is not evasion. A label that claims less is more honest when we know less, and the discomfort surgeons feel about vaguer terminology is mostly discomfort about admitting the vagueness was always there.
Man Ther 2016;23:57–68 · PMID 27083390 ·
doi:10.1016/j.math.2016.03.009 · earlier version of the argument: Br J Sports Med 2008;43(4):259–264, PMID 18838403
What the collection adds up to
Read in sequence, these entries describe a specialty that has become very good at generating evidence and is still working out what its evidence licenses. The trials are increasingly strong, the registries extraordinary, and the methodological self-audit of theme seven is more candid than most fields manage. What seems in shorter supply is philosophical work — and not only in epistemology.
Some of what is needed is analytic metaphysics. Whether fracture union is a biological state awaiting discovery or a clinical judgment imposed on a continuous process is not an empirical question, and no larger trial will settle it. Nor will one tell us whether "impingement syndrome" or "degeneration" names a natural kind or a useful grouping, or what makes a surgical technique the same technique across two surgeons and two decades. Theme ten collects the papers that come closest to asking, and it is a short theme because there are not many.
Some is meta-ethics and value theory. Every outcome measure carries a claim about what makes a life go better, and every threshold for clinical significance is a judgment about how much improvement matters and to whom. Choosing between operating and accepting is a choice over distributions of harm and benefit, which needs some account of how to weigh them. These are normative questions, and they are mostly answered implicitly — by instrument design and by convention — rather than argued out.
Much of it is epistemology proper: theory-laden observation, inference from mechanism, the calibration of confidence against accuracy, the conditions under which experience yields knowledge, and the reference class problem that arises whenever group evidence meets an individual patient — including when the group is a training set. Each tends to be settled by something other than evidence: a label whose causal claim was never tested, a prevalence never measured, a threshold set by those who developed the operation, a confidence that may be less well calibrated than it feels. None of these is irrational. Each is a reasonable local decision, and each quietly shapes what the evidence will subsequently appear to show.
If there is a practical conclusion, it is not scepticism, which is cheap and stops thought. The best entry in the last theme is a randomised trial that worked, and the collection would be dishonest without it. It is rather that these commitments are choices, that they are easy to make without noticing, and that the habit worth cultivating — in my own practice as much as anyone's — is asking what they were and what else they might have been.