Rob Parisien, MD, FAAOS
Theme One

What surgery looks like without belief

Can a procedure survive being separated from the expectation that surrounds it?

Surgery resisted placebo controls for most of a century, on grounds that were part ethical and part unexamined. The ethical objection was serious and remains unresolved — a sham operation inflicts real harm on a subject with no prospect of benefit to them, and the exchange between Macklin and Freeman in the further reading is worth going to before deciding the field was merely slow. When the controls finally arrived they did something no amount of case-series accumulation could: they separated the operation from the reference class the surgeon had assigned it to. Start here, because the base rate this theme establishes should sit underneath everything that follows.

1

Use of Placebo Controls in the Evaluation of Surgery: Systematic Review

BMJ · 2014 · Wartolowska K, Judge A, Hopewell S et al. · Systematic review, 53 trials

What it does: Collects every placebo-controlled surgical trial then published and asks a simple question of the set. Improvement occurred in the placebo arm in 74% of trials. In just over half, there was no significant difference between the real operation and the sham.

Epistemological angle: This is the base rate for surgery itself, and it belongs in a surgeon's head before any individual outcome study is read. When roughly half of the procedures that have been tested this way fail to beat placebo, an uncontrolled case series reporting good results carries almost no evidential weight — it is exactly what you would expect to see whether or not the operation does anything. The paper also answers the practical objection to sham trials by demonstration: they were done, they were feasible, and the risk was low.
My take: This is a number worth carrying around, and I did not know it for most of my career. Without some sense of how often operations beat placebo when properly tested, there is no calibrated prior to read a new trial against.
BMJ 2014;348:g3253 · PMID 24850821 · doi:10.1136/bmj.g3253
2

A Controlled Trial of Arthroscopic Surgery for Osteoarthritis of the Knee

NEJM · 2002 · Moseley JB Jr, O'Malley K, Petersen NJ et al. · Sham-controlled RCT, n=180

What it does: Randomises patients with knee osteoarthritis to arthroscopic débridement, arthroscopic lavage, or placebo skin incisions with no arthroscope inserted. At no point across 24 months did either real operation produce less pain or better function than the placebo. The confidence intervals excluded any clinically meaningful difference.

Epistemological angle: Before this trial, the evidence for knee arthroscopy in osteoarthritis was consistent, abundant, and entirely uncontrolled. Patients did improve. The inference from "they improve after surgery" to "the surgery improved them" is the one this design breaks, and it turns out to have been carrying the whole weight. Regression to the mean, natural history, and the substantial expectation effects of an operation account for the observed benefit without requiring the mechanism anyone believed in.
My take: The uncomfortable part is not that the operation failed. It is that thousands of surgeons had watched patients get better and drawn the obvious conclusion, and the obvious conclusion was wrong. Clinical experience is a real form of knowledge, but it is structurally blind to exactly this failure mode, and no amount of it accumulates into a control group.
N Engl J Med 2002;347(2):81–88 · PMID 12110735 · doi:10.1056/NEJMoa013259 · editorial: Felson & Buckwalter, NEJM 2002;347(2):132–133, PMID 12110742
3

Arthroscopic Partial Meniscectomy versus Sham Surgery for a Degenerative Meniscal Tear

NEJM · 2013 · Sihvonen R, Paavola M, Malmivaara A et al. (FIDELITY) · Double-blind sham-controlled RCT, n=146

What it does: Takes patients with a degenerative medial meniscal tear and no osteoarthritis — the classic indication — and randomises them to partial meniscectomy or a sham arthroscopy. At twelve months there was no difference on Lysholm, WOMET, or pain after exercise.

Epistemological angle: The tear was the reference class. It was the finding that assigned the patient to the operable group, that justified the arthroscope, and that explained the pain afterwards. FIDELITY shows the assignment does not license the operation: removing the thing did not remove the symptom. That is a weaker claim than showing the tear to be causally inert, and the difference is worth holding onto — a process may be irreversible by the time it is found, or the surgery's own costs may offset what it gains. But it is enough for the clinical question, because the assignment's entire purpose was to justify removal. Read alongside entry 12, which establishes that most people of this age have such a tear and most of them have no pain, the two papers form a complete argument — here is the base rate that should have warned us, and here is the trial that eventually did.
My take: The response is as instructive as the result. A formal rebuttal appeared in Arthroscopy within months arguing the trial was flawed and unrepresentative, and the practice pattern took the better part of a decade to move. That is not corruption. That is what it looks like when disconfirming evidence arrives about a procedure a specialty has organised itself around.
N Engl J Med 2013;369(26):2515–2524 · PMID 24369076 · doi:10.1056/NEJMoa1305189 · rebuttal: Elattrache N, Lattermann C, Hannon M, Cole B. Arthroscopy 2014;30(5):542–543, PMID 24642105
4

The Vertebroplasty Sequence: Two Null Trials, Then One That Was Not

NEJM · 2009 · Buchbinder R et al. (n=78) and Kallmes DF et al., INVEST (n=131), independent sham-controlled RCTs published in the same issue · and Lancet · 2016 · Clark W, Bird P, Gonski P et al., VAPOUR, double-blind placebo-controlled RCT (n=120)

What it does: Two research groups on opposite sides of the world independently randomised patients with painful osteoporotic vertebral fractures to cement injection or a sham procedure. Neither found any advantage at any time point; both arms improved substantially. Seven years later VAPOUR repeated the design in a deliberately different population — fractures under six weeks old, severe pain, a specific filling technique — and vertebroplasty won, with 44% against 21% below a pain score of 4 out of 10 at fourteen days.

Epistemological angle: Two lessons in sequence, and the second undoes the comfortable reading of the first. A single negative trial invites the reply that this trial was flawed; two independent trials in different countries with different designs and the same null move the question from the trials onto the intervention, and the epistemic gain comes from the independence rather than from the combined sample size. INVEST adds that blinding and crossover-resistance are separate problems, with 51% of its sham arm crossing to real vertebroplasty by three months against 13% the other way. Then VAPOUR states the reference class problem as a trial. It did not refute the 2009 results; it redrew the boundary of who was being asked about. "Does vertebroplasty work" was never a well-formed question — only "in fractures of what age, what severity, filled to what extent" is answerable, and every null result carries an implicit population that the headline discards. Whether VAPOUR's chosen class is the right one remains contested, largely on the adequacy of its blinding.
My take: Cement stabilises a fracture. It is difficult to think of a more mechanically obvious intervention in all of orthopaedics, and in two trials it did nothing. I would hold onto that whenever I catch myself reasoning from mechanism to benefit, which is most days. But VAPOUR is why this is one entry rather than two, and it is the more useful half. It is easy to enjoy sham trials when they confirm a suspicion that we operate too much. The discipline is in accepting the same design when it says the opposite, and in noticing that both answers are about populations rather than about procedures.
N Engl J Med 2009;361(6):557–568, PMID 19657121, doi:10.1056/NEJMoa0900429 · and 361(6):569–579, PMID 19657122, doi:10.1056/NEJMoa0900563 · editorial: Weinstein JN, N Engl J Med 2009;361(6):619–621, PMID 19657127, doi:10.1056/NEJMe0905889 · and Lancet 2016;388(10052):1408–1416, PMID 27544377, doi:10.1016/S0140-6736(16)31341-1
5

The Two Decompression Trials

Lancet · 2018 · Beard DJ, Rees JL, Cook JA et al. (CSAW, n=313) · and BMJ · 2018 · Paavola M, Malmivaara A, Taimela S et al. (FIMPACT, n=210) · Independent placebo-controlled RCTs

What it does: Both trials test arthroscopic subacromial decompression against a placebo arthroscopy in which the arthroscope is introduced and the bone and soft tissue are left alone. CSAW ran three arms — decompression, arthroscopy only, no treatment — across 32 UK hospitals and 51 surgeons, and found no difference between the two surgical groups on the Oxford Shoulder Score at six months. Both surgical arms beat no treatment by a margin the investigators judged not clinically important. FIMPACT ran the comparison in Finland with pain at rest and on activity at 24 months as co-primary outcomes and reached the same conclusion: decompression offered no benefit over diagnostic arthroscopy.

Epistemological angle: The independence argument of entry 4, repeated in a second procedure, and with a sharper design. CSAW's three-arm structure separates two questions that a two-arm sham trial has to answer together: what does the specific surgical act contribute, and what does the whole surgical episode contribute. The answer is that the act contributes nothing detectable and the episode contributes a little. That decomposition is what makes these trials the natural terminus of the arc begun at entry 8: Neer's mechanism was proposed in 1972, the diagnosis built on it, the operation built on the diagnosis, and forty-six years later the mechanism-specific step turned out to be the part that was doing nothing.
My take: Promoting these out of a footnote is a correction to an earlier version of this list, which cited them only in passing under Neer. They deserve to stand on their own, and the reason is CSAW's middle arm. A two-arm sham trial tells you the operation is no better than nothing much; a three-arm trial tells you which part of it was inert. We should be designing more trials that can answer that question, and I do not know why we design so few.
Lancet 2018;391(10118):329–338 · PMID 29169668 · doi:10.1016/S0140-6736(17)32457-1 · and BMJ 2018;362:k2860 · PMID 30026230 · doi:10.1136/bmj.k2860
6

Is the Placebo Powerless? An Analysis of Clinical Trials Comparing Placebo with No Treatment

NEJM · 2001 · Hróbjartsson A, Gøtzsche PC · Systematic review, 130 trials

What it does: Reviews trials containing both a placebo arm and an untreated arm — the only design that can measure the placebo effect itself — and finds little evidence of a large general clinical effect, with small effects on subjective outcomes such as pain. Their 2010 Cochrane update adds that physical placebos, including sham procedures, produce larger responses than pills.

Epistemological angle: This complicates the lazy reading of every trial above. "It was just the placebo effect" sounds like an explanation, but if placebo is generally weak, then a sham arm matching a real operation is a far harsher verdict than that phrase implies — it means the operation was doing very little, not that placebo was doing a great deal. And the finding that procedural placebos outperform pill placebos is what makes surgical sham controls necessary rather than borrowable: surgery could not import its intuitions from drug trials, because the placebo in surgery is a different and larger thing.
My take: This is the entry most likely to be skipped and it is the one that makes the rest rigorous. Without it, "sham was as good as surgery" gets absorbed as a comforting story about the power of the mind, when the harder reading is that the operation was close to inert.
N Engl J Med 2001;344(21):1594–1602 · PMID 11372012 · doi:10.1056/NEJM200105243442106 · update: Cochrane Database Syst Rev 2010;(1):CD003974, PMID 20091554
7

Vertebral Augmentation After Recent Randomized Controlled Trials: A New Rise in Kyphoplasty Volumes

JACR · 2015 · Cox M, Levin DC, Parker L, Morrison W, Long S, Rao VM · Medicare billing analysis, 2006–2013

What it does: Tracks procedure volumes after the 2009 sham-controlled vertebroplasty trials. Vertebroplasty fell sharply. Kyphoplasty — a mechanistically similar procedure never subjected to the same test — dipped and then rose, partly offsetting the decline in total vertebral augmentation.

Epistemological angle: The cleanest empirical demonstration of the adopt-versus-abandon asymmetry available in orthopaedics. Kyphoplasty was never held to the standard eventually applied to vertebroplasty; it entered practice on plausibility and stayed there. When disconfirming evidence arrived for the tested procedure, demand migrated to the untested cousin. No individual made an irrational decision, and the aggregate result was that a sham-controlled refutation produced a change in billing codes rather than a change in whether patients had cement injected.
My take: This is what makes me sceptical that better evidence alone fixes anything. The evidence arrived, it was excellent, it was widely publicised, and the practice found a route around it that required nobody to disagree with it.
J Am Coll Radiol 2015;13(1):28–32 · PMID 26546300 · doi:10.1016/j.jacr.2015.08.025

Also worth reading

Macklin R versus Freeman TB, Vawter DE, Leaverton PE et al. Is sham surgery permissible? Opposing essays, N Engl J Med 1999;341(13):992–996, PMID 10498498, doi:10.1056/NEJM199909233411312 · and 341(13):988–992, PMID 10498497, doi:10.1056/NEJM199909233411311. The ethical cost of the design that makes this theme possible, argued by two competent parties neither of whom is wrong. Developed further in Horng S, Miller FG, N Engl J Med 2002;347(2):137–139, PMID 12110744, doi:10.1056/NEJMsb010576.