Rob Parisien, MD, FAAOS
Research → The Canon

The canonical epistemology-oriented literature in orthopaedics

Exemplar papers — orthopaedic, with selected work from general medicine and philosophy — grouped into ten themes · compiled July 2026

A catalogue of exemplar papers, each selected to illustrate a specific epistemological problem in orthopaedic practice. The organising question is not does this operation work, but the one that precedes it: how do we know what we claim to know, and what would have to be true for this evidence to bear on this patient?

Five difficulties recur, and none of them is anyone's fault. The first is prior to the rest: a great deal of what we count, analyse and report is a category imposed on something continuous, and the placement of the boundary does work that the data are then held to have done. Given a category, four further problems follow. Observation is theory-laden: an imaging finding is not raw data but something read under a framework that shapes what counts as pathology. Causal inference is often carried by assumption — from mechanism to benefit, from post-operative improvement to operative effect. Confidence and knowledge are easily confounded, since confidence grows with experience whether or not accuracy does. And success is defined by whoever chose the outcome measure, which is a normative judgment as much as a methodological one.

Three of these turn, in the end, on the reference class problem — theory-ladenness and the choice of outcome measure are genuinely separate difficulties, the second of them normative rather than epistemic. Every piece of evidence we use is a statement about a group, and none of them is the patient in front of us. Using any of it requires placing that patient into a class, and the evidence cannot make the placement for us. Hájek's point, given its own entry at 17, is that this is no frequentist quirk that Bayesians escape — choosing a prior is choosing a reference class. It reappears at the end of the collection in mechanical form, since a trained model's reference class is its training distribution and nothing in the model selects it.

Papers were selected on one criterion: that what they teach is about inference rather than about any particular joint or procedure. Some are exemplary, some cautionary, and the most useful are both. Where a study serves as a cautionary case, the difficulty is treated as the field's rather than the authors' — partly because that is usually the truth of it, and partly because an error a careful person makes teaches more than a paper to be dismissed. That policy is stated once here rather than repeated at each entry.

Ten themes

A note on method

Every citation was checked against PubMed rather than written from memory: author, year, journal, volume, pages, PMID and DOI. Every paper named anywhere in this collection carries its own PMID and a link to its DOI, including papers cited alongside a main entry. Where PubMed carries no DOI — several of the older papers — that is stated rather than a plausible string supplied. Where an entry lies outside PubMed's scope, that is flagged in its citation line.

Each theme closes with a short further-reading list. Those papers are not lesser work; in most cases they make a point a main entry already makes, and the two-tier arrangement exists so that the reading burden stays manageable without anything being thrown away. Expect traffic in both directions as the collection grows.

Selection was not blind to how much a paper was argued over. A study that drew twenty letters often teaches better than a quietly correct one, because the argument is part of the lesson. Where a published rebuttal exists, it is cited alongside, and readers are encouraged to go to it.

Five entries are my own work or my group's. Each is criticised in its amber box, sometimes at more length than the others receive, but a reader should discount them accordingly and I would rather say so here than at each one.

The blue passages are analysis. The amber passages are my own opinion, offered tentatively and open to correction — I have no doubt some of them are wrong.

One gap worth naming

No orthopaedic study of citation distortion. The general phenomenon is described elsewhere in medicine: claims propagate through citation networks, gaining apparent authority with each hop while the evidence beneath them stays where it was, and disconfirming studies are cited less often than confirming ones. Entry 15 is the closest thing to it, and it is a study of general medicine rather than of us — highly cited findings contradicted or attenuated at a rate of roughly one in three, with the damage concentrated in small and non-randomised work. That is the study I would like to see run on the orthopaedic literature. Given how much of this collection concerns claims that outlived their support, the mechanism by which they survived seems the natural next thing to examine.

A second gap named earlier — the near-absence of Bayesian reanalysis with explicit priors — is left standing but qualified. Searching across the major fracture and arthroplasty trials, I found one clean example, entry 21. It is not that nobody has thought to do it; work of this kind has been written and submitted, including my own. What the literature may be missing, it is missing nearer the point of publication than the point of conception.

What the collection adds up to

Read in sequence, these entries describe a specialty that has become very good at generating evidence and is still working out what its evidence licenses. The trials are increasingly strong, the registries extraordinary, and the methodological self-audit of theme seven is more candid than most fields manage. What seems in shorter supply is philosophical work — and not only in epistemology.

Some of what is needed is analytic metaphysics. Whether fracture union is a biological state awaiting discovery or a clinical judgment imposed on a continuous process is not an empirical question, and no larger trial will settle it. Nor will one tell us whether "impingement syndrome" or "degeneration" names a natural kind or a useful grouping, or what makes a surgical technique the same technique across two surgeons and two decades. Theme ten collects the papers that come closest to asking, and it is a short theme because there are not many.

Some is meta-ethics and value theory. Every outcome measure carries a claim about what makes a life go better, and every threshold for clinical significance is a judgment about how much improvement matters and to whom. Choosing between operating and accepting is a choice over distributions of harm and benefit, which needs some account of how to weigh them. These are normative questions, and they are mostly answered implicitly — by instrument design and by convention — rather than argued out.

Much of it is epistemology proper: theory-laden observation, inference from mechanism, the calibration of confidence against accuracy, the conditions under which experience yields knowledge, and the reference class problem that arises whenever group evidence meets an individual patient — including when the group is a training set. Each tends to be settled by something other than evidence: a label whose causal claim was never tested, a prevalence never measured, a threshold set by those who developed the operation, a confidence that may be less well calibrated than it feels. None of these is irrational. Each is a reasonable local decision, and each quietly shapes what the evidence will subsequently appear to show.

If there is a practical conclusion, it is not scepticism, which is cheap and stops thought. The best entry in the last theme is a randomised trial that worked, and the collection would be dishonest without it. It is rather that these commitments are choices, that they are easy to make without noticing, and that the habit worth cultivating — in my own practice as much as anyone's — is asking what they were and what else they might have been.