Theme Eight
The surgeon as instrument
What happens to evidence when the treatment is a human being?
A drug is the same molecule in every trial. A surgical procedure is a person, on a particular day, at a particular point on their learning curve, classifying the pathology by eye. Everything orthopaedics borrowed from pharmaceutical trial methodology inherited an assumption that does not hold here. This theme is about the instrument: how well it is calibrated, how it perceives, the trial designs built around its variability, and what it turns out the operation was doing when it worked.
36
Do Orthopaedic Surgeons Acknowledge Uncertainty?
CORR · 2016 · Teunis T, Janssen S, Guitton TG, Ring D, Parisien R · Cross-sectional survey, n=242
What it does: Measures surgeons' recognition of uncertainty and looks for what predicts it. Recognition was uniformly low and did not improve with years in practice, while confidence bias increased with experience. Better statistical understanding was the strongest independent predictor of acknowledging uncertainty. Greater trust in the orthopaedic evidence base was independently associated with acknowledging less.
Epistemological angle: That last association is the finding. Surgeons who believe orthopaedics rests on solid evidence are less able to say "I don't know" — which is exactly backwards, given that the honest reading of the preceding thirty-four entries is that the evidence base does not support such confidence. The paper documents a specialty whose self-assessment runs contrary to its actual epistemic position. That is not a knowledge deficit; it is a calibration failure, and calibration failures are by their nature invisible from the inside. Read with entry
33, which suggests the acknowledgement may exist and be getting stripped out at the point of publication rather than never forming.
My take: Three criticisms, and they strengthen it. The instrument measures self-reported acknowledgement rather than actual calibration — a surgeon may score well and still judge badly. The design is cross-sectional, so rising confidence with experience is equally explicable as a cohort effect. And the sample came from the Science of Variation Group, meaning it drew on the most uncertainty-aware population in the specialty, and recognition was still uniformly low. Every limitation points the same way as the conclusion. I wrote this one, and I think the real number is worse than we reported.
Clin Orthop Relat Res 2016;474(6):1360–1369 · PMID 26552806 ·
doi:10.1007/s11999-015-4623-0 · erratum 474(6):1530–1531, PMID 26861152; Editor's Spotlight, Leopold SS, 474(6):1356–1359, PMID 26818597
37
Two Proposals for Trials Whose Instrument Is a Person: IDEAL and Expertise-Based Randomisation
Lancet · 2009 · McCulloch P, Altman DG, Campbell WB et al., consensus framework · and BMJ · 2005 · Devereaux PJ, Bhandari M, Clarke M, Montori VM, Cook DJ, Yusuf S, Sackett DL et al., methodological proposal
What it does: IDEAL proposes a five-stage framework for evaluating surgical innovation — idea, development, exploration, assessment, long-term study — built explicitly around confounding by operator, team, learning curve and perceived equipoise. Devereaux and colleagues propose a change to allocation itself: randomise patients to a surgeon expert in procedure A or a surgeon expert in procedure B, rather than to a procedure that whichever surgeon is available then performs.
Epistemological angle: Two responses to the same structural problem, and they are complementary rather than alternative. IDEAL states the problem: a surgical treatment is not a stable standardisable input the way a drug is, so a trial design imported wholesale from pharmacology answers a question about an intervention that does not exist in the form assumed. The learning curve alone means a procedure evaluated early is a different intervention from the same procedure evaluated late, and conventional analysis treats them as one. Expertise-based randomisation is the design consequence: if the same surgeon delivers both arms and is better at one of them, the trial measures A against B as performed by people trained mostly in A. Making the surgeon part of the allocated intervention is the right response to a variable that is not noise but structure — at the cost, which the authors accept, that the result then attaches to a procedure-plus-practitioner package rather than to a procedure.
My take: Both are widely cited and neither is much followed. Orthopaedic innovation still arrives fully formed in a case series, which says something about incentives rather than about the frameworks. The expertise-based design has an obstacle worth naming honestly, because it is not logistical: it requires a surgeon to concede in advance and in writing that a colleague is better at something. That is a real cost, and pretending the barrier is purely practical has not helped anyone adopt it in twenty years.
38
The Invisible Gorilla Strikes Again: Sustained Inattentional Blindness in Expert Observers
Psychol Sci · 2013 · Drew T, Võ ML-H, Wolfe JM · Experimental study, 24 radiologists
What it does: Asks radiologists to perform a familiar lung-nodule detection task on chest CT. An image of a gorilla, forty-eight times the size of the average nodule, was inserted into the final case. Eighty-three per cent of the radiologists did not report it. Eye tracking showed that most of those who missed it had looked directly at it.
Epistemological angle: The front matter of this collection asserts that observation is theory-laden. This is the entry that demonstrates it, and it demonstrates something stronger than the usual formulation. The claim is not merely that expectation colours interpretation of what was seen; it is that expertise organises the visual search itself, so that the framework determines what enters awareness at all. The gorilla was fixated and not seen. That places the effect upstream of judgment, where no amount of care in reasoning can reach it, and it is why entry
49's finding about classification disagreement should not be read as carelessness. A skilled observer is a tuned instrument, and tuning is subtraction as well as amplification.
My take: This is the one I like to point out, because it is the only entry here that people find funny, and the laugh is the point at which the argument lands. The thing worth taking away is that expert perception is constructed, not that any particular specialty is careless. It also cuts against a comfort I notice in myself: I read a scan differently having formed an impression from the history, and I have always filed that under experience rather than under the thing this paper measures.
39
Outcomes After Rotator Cuff Repair Are Not Related to Structural Healing
Arthroscopy · 2021 · Holtedahl R, Bøe B, Brox JI · Meta-regression, 64 RCT and 19 cohort arms
What it does: Pools trial arms of rotator cuff repair. Retear rate was around 20% at a median of thirteen months. Age and tear size predicted retear. Clinical outcome did not track whether the repair had actually healed.
Epistemological angle: This has something close to the shape of a Gettier case, and it is worth being careful about how close. The belief "rotator cuff repair helps patients" is true. The justification every surgeon holds — the torn tendon is reattached, it heals, function follows — is disconnected from the truth-maker in a substantial fraction of the people who benefit. Whether that is Gettier proper or simply a true belief resting on a false explanation is arguable, and the arguable part is instructive rather than a defect. Either way the practical upshot holds: the reason the belief is true is not the reason we have for holding it. The same structure applies to Moseley and vertebroplasty, where patients genuinely improved and the stated mechanism was doing none of the work.
My take: If a fifth of repairs fail structurally and those patients do as well, then something other than the repair is producing the benefit and we do not know what it is. That is a more interesting research question than another comparison of fixation constructs.