Rob Parisien, MD, FAAOS
Theme Ten

Forced categories

What happens when a continuous or subjective phenomenon is made to be one thing or another?

This theme is last because it is first. Every preceding theme presupposes that there is a kind here to have base rates about, to randomise treatments for, to measure with surrogates and to argue over. Where the category does not carve — where union, impingement, degeneration or clinically meaningful improvement is a boundary drawn across a continuum rather than a joint the world already has — nothing downstream repairs it. You cannot compute a prevalence for a thing that is not there to be prevalent, and no larger trial will settle whether it is. These are questions of category, vagueness and identity, and they are not empirical. It is also the least developed area in this collection, which is why the marginal return on work here seems to me higher than anywhere else on the list.

47

The Cost of Dichotomising Continuous Variables

BMJ · 2006 · Altman DG, Royston P · Statistical note

What it does: One page, setting out what is lost when a continuous measurement is converted into two categories. Information is discarded; a step change in risk is assumed where the data show a gradient; individuals close to the cutpoint on opposite sides are treated as different in kind while those at opposite ends of the same category are treated as the same; statistical power falls; and where the cutpoint is chosen after inspecting the data, the resulting effect estimate is optimistic.

Epistemological angle: The statistical statement of this theme's thesis, and deliberately placed first because it is the cheapest to grasp and the hardest to unsee. Once the argument is in view it applies to almost every categorical variable in orthopaedic research. Union and non-union is a dichotomy imposed on a continuous process of mineralisation. Success and failure on a patient-reported outcome is a dichotomy imposed on a continuous distribution of experience. Displaced and undisplaced is a dichotomy imposed on a continuum of millimetres. In each case the categorical variable is the thing we count, analyse and report, and Altman and Royston's point is that the choice of where to cut is doing work that the data are then held to have done.
My take: The last of their objections is the one I would emphasise for our field. A cutpoint chosen after the results are known is not a measurement decision but an inference dressed as one, and radiographic union thresholds have largely been arrived at that way. This is a one-page paper and I have never seen it cited in an orthopaedic manuscript.
BMJ 2006;332(7549):1080 · PMID 16675816 · doi:10.1136/bmj.332.7549.1080
48

Guidance for Modifying the Definition of Diseases: A Checklist

JAMA Intern Med · 2017 · Doust J, Vandvik PO, Qaseem A et al. · Consensus checklist, thirteen-member multidisciplinary group

What it does: Observes that guideline panels frequently widen disease definitions, increasing the proportion of a population labelled unwell, and that no guidance existed for panels considering such a change. Proposes an eight-item checklist covering the number of people affected by the change, the trigger for it, the prognostic ability and precision of the revised definition, and the balance of benefits and harms — with the explicit requirement that a definitional change be justified rather than announced.

Epistemological angle: The normative counterpart to entry 47. Altman and Royston show what dichotomising costs; this treats the placement of the cut as a decision requiring argument, which is precisely the move the orthopaedic literature has not made. The orthopaedic instance is the osteoporosis T-score: a threshold at 2.5 standard deviations below a young-adult reference mean, applied to a continuous and unimodal distribution of bone density, which brought a disease into existence by stipulation and set the denominator for everything subsequently claimed about treating it. Note also what the checklist implies about entries 8 to 10 — if a definitional change alters how many people are called unwell, then the label's effects documented there are a consequence of a decision somebody made, not a fact about the world.
My take: I would like to see a version of this checklist applied retrospectively to the definitions we use every day, beginning with union. My expectation is that most of them would fail item four — precision and accuracy of the definition — not because they are poorly made but because nobody was ever required to state what the definition was for.
JAMA Intern Med 2017;177(7):1020–1025 · PMID 28505266 · doi:10.1001/jamainternmed.2017.1302
49

The Neer Classification System for Proximal Humeral Fractures: Interobserver Reliability and Intraobserver Reproducibility

JBJS Am · 1993 · Sidor ML, Zuckerman JD, Lyon T, Koval K, Cuomo F, Schoenberg N · Observer-agreement study, 50 fractures, 5 observers

What it does: Has five observers of varying seniority classify the same fifty fractures twice, six months apart. Interobserver kappa 0.48 to 0.52. All five agreed on roughly 30% of fractures. Collapsing the system from sixteen categories to six did not improve agreement.

Epistemological angle: A classification can look like objective anatomical description while functioning as a low-reliability perceptual judgment. Every study that stratifies treatment or outcome by Neer type silently inherits this noise, which means a substantial literature is built on a variable that five experts agree about less than a third of the time. That simplification did not help is the diagnostic detail: the problem is in the seeing, not the taxonomy.
My take: A parallel from the hip is worth putting beside this. Frandsen's 1988 work on the Garden classification found agreement in 22 of 100 fractures, and the field responded by collapsing Garden into displaced versus undisplaced. Coarsening did not rescue Neer, so it is not a general remedy — but where the coarser distinction is the one that actually drives treatment, it is honest about how much the observer can see. That is one of the few clean self-corrections in this collection.
J Bone Joint Surg Am 1993;75(12):1745–1750 · PMID 8258543 · doi:10.2106/00004623-199312000-00002 · Garden: Frandsen PA et al., JBJS Br 1988;70(4):588–590, PMID 3403602
50

Development of the Radiographic Union Score for Tibial Fractures (RUST)

J Trauma · 2010 · Whelan DB, Bhandari M, Stephen D, Kreder H, McKee MD, Zdero R, Schemitsch EH · Reliability study, 45 radiograph sets, 7 reviewers

What it does: Establishes that a radiographic scoring system for tibial fracture healing is reproducible — interobserver ICC 0.86, intraobserver 0.88. The authors state plainly that no gold standard for union exists against which the score has been validated.

Epistemological angle: The cleanest available separation of reliability from validity. Seven observers agreeing tells you the instrument is consistent; it tells you nothing about whether it measures a healed, painless, usable limb. RUST has since become the most widely adopted union scale in trauma research and functions as a de facto trial endpoint. The caveat the authors printed has not travelled with it.
My take: This is not a criticism of the paper, which did exactly what it claimed and flagged what it had not done. It is a criticism of what a field does with a well-made instrument: reliability is easy to demonstrate and validity is hard, so we demonstrate reliability and then quietly start behaving as though we had shown the other thing.
J Trauma 2010;68(3):629–632 · PMID 19996801 · doi:10.1097/TA.0b013e3181a7c16d
51

Rotator Cuff Related Shoulder Pain: Assessment, Management and Uncertainties

Man Ther · 2016 · Lewis J · Position paper

What it does: Argues that "subacromial impingement syndrome," "rotator cuff tendinopathy" and the tear categories should be collapsed into a single umbrella term that claims less — rotator cuff related shoulder pain — on the grounds that the existing labels assert a causal mechanism that has not been shown and that imaging findings correlate poorly with symptoms.

Epistemological angle: The only entry here that engages the nosological question directly: not "does this treatment work" but "does this named entity correspond to anything real enough to be a diagnosis." Lewis had made the same argument eight years earlier and the field had not moved; taken together the two papers show a specialty revising a category in real time, which is a slow and visible process rarely captured in the literature.
My take: Renaming is not evasion. A label that claims less is more honest when we know less, and the discomfort surgeons feel about vaguer terminology is mostly discomfort about admitting the vagueness was always there.
Man Ther 2016;23:57–68 · PMID 27083390 · doi:10.1016/j.math.2016.03.009 · earlier version of the argument: Br J Sports Med 2008;43(4):259–264, PMID 18838403