Theme Four
Bayes, the prior, and the reference class
What did the p-value throw away, and what would a prior have kept?
Theme three establishes that base rates exist and are often extreme. This theme is about what happens next. It opens with four papers that state the problem in general form — the base rate of true findings in a literature, how often established practice reverses when tested, the question a p-value does not answer, and the identity between a prior and a reference class — and then turns to whether clinicians in fact reason this way. Two entries measure surgeons directly, and the results are not flattering to any of us.
14
Why Most Published Research Findings Are False
PLoS Med · 2005 · Ioannidis JPA · Analytical essay with simulations · general science
What it does: Models the probability that a published research claim is true as a function of study power, bias, the number of teams working the same question, and — the load-bearing term — the ratio of true to null relationships among those a field chooses to test. Under most realistic combinations of these, the post-study probability that a claimed positive finding is true falls below one half.
Epistemological angle: This is theme three's argument turned on the literature rather than the patient. A published positive finding is a test result. Interpreting it requires a prior, and the prior is the proportion of hypotheses in that field that are true — a quantity nobody measures and everybody implicitly sets at something close to one. Ioannidis makes the reference class of a research claim explicit: it belongs to the set of studies of its design, in its field, at its power, and its credibility is a property of that set before it is a property of the paper. His last observation is the sharpest, and it is a claim about theme nine as much as this one: where prejudice and financial interest are strong, a published finding may be an accurate measurement of the field's prevailing bias rather than of the world.
My take: The title is doing rhetorical work and the paper is more careful than the title. It is a conditional argument, not a census, and how much of it applies to orthopaedic surgical trials depends on parameters we have never estimated for ourselves. I would rather read it as a framework we have not yet run our own numbers through than as a verdict already delivered on us. That said, the exercise of guessing at those parameters for our specialty is uncomfortable in a way I think is instructive.
15
The Reversal Rate: How Often Established Practice Fails When Tested
Mayo Clin Proc · 2013 · Prasad V, Vandross A, Toomey C et al. · 2,044 original articles over ten years · and JAMA · 2005 · Ioannidis JPA · 49 highly cited clinical studies
What it does: Prasad classifies every original article published over a decade in one high-impact journal by what it did to practice. Of 363 articles testing an established practice, 146 (40.2%) reversed it and 138 (38.0%) reaffirmed it. Ioannidis asks the same question of celebrated findings rather than of standard care: of 49 highly cited clinical studies, 45 claimed an intervention effective, and of those, seven were later contradicted outright and seven more had found effects stronger than subsequent work. Twenty replicated. The breakdown by design is the sharp part — five of six highly cited non-randomised studies were contradicted or attenuated, against nine of 39 randomised trials (p=0.008), and among the trials the vulnerable ones were the small ones.
Epistemological angle: Together these supply the prior that entry
14 models theoretically. When an established practice is finally subjected to a proper test it fails roughly as often as it survives, and a celebrated finding has about a one-in-three chance of shrinking or reversing. That number should govern how much evidential weight "this is standard of care" and "this is the landmark trial" carry, and in practice both carry far more. Prasad's denominators add a second finding that is easy to miss: 73% of the literature examined new practices and 27% examined existing ones. The field spends roughly three times the effort finding things to start as finding things to stop, and then treats the untested incumbent as the safe default.
My take: Between them these are the most useful numbers in the collection for a surgeon's own calibration, and both sit in general medicine rather than in ours. Nobody has run either analysis on the orthopaedic literature. That would be a straightforward and worthwhile study, and my guess is that our rate is not better. I would rather be shown wrong by data than continue guessing.
16
Toward Evidence-Based Medical Statistics: The P Value Fallacy and The Bayes Factor
Ann Intern Med · 1999 · Goodman SN · Two-part methodological essay · general medicine
What it does: Part 1 argues that the standard hypothesis-testing apparatus is an amalgam of two incompatible statistical philosophies, and that the p-value fallacy is the belief that a single number can carry both the long-run error behaviour of a procedure and the evidential meaning of the particular result in front of you. Part 2 proposes the Bayes factor as the measure that separates them, and shows that p-values systematically overstate the evidence against the null.
Epistemological angle: Most of this collection asserts at some point that the p-value answers a question the surgeon did not ask. This is where that claim is actually argued, and the argument is more interesting than the slogan. Goodman's point is not that frequentist statistics are wrong but that they are answering a question about procedures — how often would this method mislead me in the long run — and that clinicians read the answer as though it were about this result, this patient population, this decision. The Bayes factor is offered as the piece that makes the difference visible: evidential strength first, and then, separately and explicitly, whatever background knowledge you bring to it. Read this before entry
21, which is the same argument carried out on real orthopaedic data.
My take: Part 1 is the one to read if you read only one, and it is twenty-five years old, which should give us pause. The critique was available, clear, published in a journal every academic clinician reads, and the field's statistical practice is essentially unchanged. That is a fact about incentives and training rather than about the argument's quality, and it is the kind of fact this collection keeps running into.
17
The Reference Class Problem Is Your Problem Too
Synthese · 2007 · Hájek A · Philosophical paper · outside PubMed's scope; citation verified against the journal record
What it does: Argues that the reference class problem — that any individual belongs to indefinitely many groups with different frequencies, and nothing in the evidence selects among them — is not a local defect of frequentism to be escaped by adopting some other interpretation of probability. Subjectivism, logical probability and propensity accounts each face a structurally identical problem at some point in their machinery. For the Bayesian, it surfaces in the choice of prior.
Epistemological angle: This is the beam the whole collection rests on, and it belonged in the list from the start rather than only in the introduction. Every entry in theme three is a reference class problem; so is the disagreement in entry
24, the threshold in entry
25, and the population question in entry
4. Hájek's contribution is to close the exit. The comfortable reading of theme four is that if surgeons would only adopt priors, the arbitrariness would go away. It does not go away; it moves. What Bayesian practice buys is not the elimination of the choice but its visibility — the prior has to be written down, which means it can be argued with, varied, and shown to matter. That is a real gain, and it is a smaller one than it is usually sold as. The point returns in mechanical form at entry
42: a trained model's reference class is whatever its training data made salient, the model does not choose it and cannot report it, and the question of which class a system was right about turns out to be the only one that governs whether its accuracy transfers to your patient.
My take: I find this the most clarifying paper on the list and the least likely to be read from it. The reason to include it anyway is that without it the Bayesian material reads as a technical fix for a technical problem, when the actual situation is that the judgment cannot be eliminated from anywhere and the best we can do is put it where people can see it. That is a modest conclusion and I think it is the true one.
18
Base-Rate Neglect Across Thirty-Six Years
NEJM 1978, Casscells W, Schoenberger A, Graboys TB · and JAMA Intern Med 2014, Manrai AK, Bhatia G, Strymish J, Kohane IS, Jain SH · general medicine
What it does: Casscells posed a problem to physicians, residents and students: a disease with prevalence one in a thousand, a test with a 5% false-positive rate, a positive result — what is the probability the patient has the disease? The modal answer was 95%. The answer is about 5%. Manrai re-ran it with modern clinicians thirty-six years later and got the same failure.
Epistemological angle: Taught as a pair because the replication is the finding. A single 1978 result invites the reply that training has improved since. It has not. The error is stable across four decades of curriculum reform, which suggests it is not a gap in what clinicians were taught but a feature of how the reasoning system works under this particular load. What gets neglected is the denominator, and the denominator is the reference class.
My take: The specific shape of the error is worth noting: people substitute the test's false-positive rate for the probability of disease. They are not calculating badly, they are answering a different and easier question without noticing the swap.
19
Cognitive Biases in Orthopaedic Surgery
JAAOS · 2021 · Janssen SJ, Teunis T, Ring D, Parisien R · Vignette survey, n=196 orthopaedic surgeons
What it does: Puts the Casscells problem and several others directly to orthopaedic surgeons. Across three base-rate vignettes, 43%, 88% and 35% chose answers consistent with base-rate neglect. Confirmation bias appeared in 11% to 51% depending on the scenario. Randomising the anchor moved responses by 22 percentage points; changing the frame moved them by 16.
Epistemological angle: This is the entry that brings the general literature home. Everything in entry
18 was demonstrated in physicians broadly; this demonstrates it in our specialty, with our vignettes, at rates that make it impossible to treat as somebody else's problem. The anchoring and framing results are arguably worse than the base-rate ones, because they show the same surgeon gives different answers to the same clinical question depending on an arbitrary number they were shown first.
My take: The 88% vignette gave me pause. This is not a minority struggling with a hard question but most of a thoughtful sample, which suggests the question is genuinely hard. The framing effect adds something worth sitting with: some part of surgeon-to-surgeon variation may reflect different anchors rather than substantive disagreement.
20
Musculoskeletal Surgeons Use Mixed Reasoning Rather than Pure Bayesian Strategies in Clinical Practice
PLoS One · 2026 · Parisien R, Drost A, Razi A, Ramtin S, Ring D, Janssen SJ · Survey, n=153 Science of Variation Group
What it does: Presents surgeons with eight scenarios involving test and treatment decisions under uncertainty, scoring each response on a four-point scale from non-Bayesian to fully Bayesian. Mean score 3.0. Fully non-Bayesian reasoning was rare at 8.6%; fully Bayesian reasoning accounted for 29%. Most responses were mixed — prior probability acknowledged but underweighted, without explicit updating. 85% of surgeons reasoned fully Bayesianly at least once; 42% reasoned non-Bayesianly at least once.
Epistemological angle: The interesting result is the inconsistency, not the average. The same surgeon reasons well in one scenario and badly in another, which means Bayesian competence is not a stable trait being measured. The Cronbach alpha of 0.43 makes this explicit and the authors say so: the scenarios were not tapping one underlying construct. That is a more useful finding than a clean score would have been, because it relocates the problem from the person to the situation — and situations can be redesigned in a way that dispositions cannot.
My take: A low alpha is usually reported apologetically. Here it is the point: "does this surgeon reason Bayesianly" turns out not to be a well-formed question, which is itself an argument for decision aids over exhortation. One limitation cuts in a particular direction: the Science of Variation Group is self-selected and unusually reflective, so the wider picture may well be less favourable.
21
High-Dose Dual Antibiotic Cement for Hip Hemiarthroplasty: A Post-Hoc Bayesian Analysis
Bone & Joint Journal · 2025 · Farrow L, Hudson J, George A, Reed MR, Campbell MK · Bayesian reanalysis of the WHiTE 8 RCT, n=4,406
What it does: Takes a large randomised trial whose frequentist analysis was null — odds ratio 1.44, 95% confidence interval 0.88 to 2.37, filed as no benefit — and reanalyses it with explicit priors. The posterior probability of at least some benefit from dual-antibiotic cement was 92% to 98%, depending on which prior was used.
Epistemological angle: The best demonstration in the orthopaedic literature of what the frequentist convention discards, precisely because the contrast is within a single dataset rather than between studies. The p-value answered whether the data were surprising under a null nobody believed. The posterior answers how likely the treatment is to help, which is the question a surgeon actually has. Note also that the answer depends on the prior — that is not a weakness but the honest exposure of an assumption that a frequentist analysis makes silently, and entry
17 is the reason it cannot be exposed any further than that.
My take: "No significant difference" was never a finding about the world; it was a statement about a threshold. A trial of four thousand patients told us something quite specific about infection risk and the conventional analysis reported it as nothing.
22
Does the Surgical Innovation Evidence Apply to You? How Surgeon Volume and Experience Determine Whether Adoption Helps or Harms
Clin Orthop Relat Res · 2026 · Parisien R · Volume-stratified break-even decision model
What it does: Builds a decision model for whether adopting a new surgical approach helps or harms a given surgeon, using direct anterior versus posterior hemiarthroplasty as the worked case. Three parameters carry it: the number needed to treat from the trial evidence, a learning-curve decay constant governing how quickly early excess harm resolves, and a weight expressing how much a harm counts against a benefit. Break-even is reached at different annual volumes under different assumptions, and below that volume adoption is expected to harm. An interactive version lets the reader enter their own figures.
Epistemological angle: Nearly all discussion of the reference class problem in medicine concerns the patient — which group does this person belong to, and does the trial's population contain them. This applies the same problem to the operator, which is an unusual move and, in surgery, the more consequential one. A trial's effect estimate is an average over that trial's surgeons, who were selected for competence with the technique under test; the surgeon deciding whether to adopt it is very often not a member of that class, and no feature of the published result tells them so. Where entry
37 argues that the instrument's variability should be built into allocation, this argues that it should be built into the reading of a result already published, and makes the class membership a quantity rather than a caveat. It is also the only entry here where the reference-class choice is operationalised rather than described: Hájek's point at entry
17 is that the choice cannot be eliminated, only made visible, and a model with the class as an explicit parameter is one way of making it visible enough to argue with.
My take: I wrote this, so the criticisms should come from me. The decay constant and the harm weight are the whole model, and neither has been measured — they are elicited, and a conclusion sensitive to two unmeasured parameters is an argument structure rather than a finding. It is illustrated on a single procedure pair, which is narrower than the claim it is making. And the deeper problem is one the paper concedes too quietly: it individuates surgeons by annual volume because volume is what registries record, not because volume is the variable that carves. Case mix, assistant quality and institutional support may matter more, and if they do, the reference class problem has simply reappeared inside the proposed solution. That is Hájek's point arriving at my own door, and I do not think a model can escape it — only make the choice legible and let the reader disagree with it. What I would defend is the direction: we have spent a great deal of effort asking whether a patient resembles a trial population and almost none asking whether the surgeon does.
Clin Orthop Relat Res 2026 · in press; volume, pages, PMID and DOI to be supplied at publication · interactive model at
rcpmodel.netlify.app
Also worth reading
Gill CJ, Sabin L, Schmid CH. Why clinicians are natural Bayesians. BMJ 2005;330(7499):1080–1083 · PMID 15879401 · doi:10.1136/bmj.330.7499.1080. Worth reading before the survey entries because it stops them landing as an insult. Nobody reasons from a blank prior; the question is whether the prior in use is the one you would endorse if you looked at it.
The Thessaly test, rise and collapse. Karachalios T et al., J Bone Joint Surg Am 2005;87(5):955–962, PMID 15866956, doi:10.2106/JBJS.D.02338 · Konan S et al., Knee Surg Sports Traumatol Arthrosc 2009;17(7):806–811, PMID 19399477, doi:10.1007/s00167-009-0803-3 · Goossens P et al., J Orthop Sports Phys Ther 2015;45(1):18–24, PMID 25420009, doi:10.2519/jospt.2015.5215 · Blyth M et al., Health Technol Assess 2015;19(62):1–62, PMID 26243431, doi:10.3310/hta19620. A single-centre accuracy claim of 94–96% collapsing to 54–59% as replication became independent — entry 15's general finding as one specific story. Demoted only because its base-rate lesson duplicates theme three; it remains the best teaching arc in the collection and belongs in any course built from this list.