Rob Parisien, MD, FAAOS
Theme Five

Proxies, surrogates, and moving the goalposts

Who decided what counts as success, and what did they have at stake?

Every surrogate endpoint is a claim that one reference class stands in for another — that radiographic union means a usable limb, that asymptomatic clot means fatal embolism, that a score above some threshold means a patient who is glad they had the operation. The formal condition under which such a claim holds is that the intervention's effect on the real outcome be fully mediated by the surrogate; correlation between the two is not enough and never was. The entries here are about what happens when nobody checks, and about the special case where a field sets its own passing mark.

23

The Cardiac Arrhythmia Suppression Trial (CAST)

NEJM · 1989 and 1991 · CAST Investigators; Echt DS, Liebson PR, Mitchell LB et al. · Placebo-controlled RCT, n=1,498 in the encainide and flecainide arms

What it does: Tests the hypothesis that suppressing ventricular ectopy after myocardial infarction reduces sudden death. Patients whose ectopy was successfully suppressed by encainide or flecainide were randomised to continue the drug or to placebo. The encainide and flecainide arms were stopped early: 43 arrhythmic deaths on drug against 16 on placebo, and total mortality 7.7% against 3.0% in the preliminary report. The drugs did exactly what they were designed to do to the surrogate, and killed people.

Epistemological angle: The example the surrogate literature is built on, given its own entry here because the structure deserves examination rather than citation. Every element of the reasoning was sound in isolation. Ectopy predicts sudden death — true, and well established. These drugs suppress ectopy — true, and confirmed in every patient before randomisation. Therefore suppressing ectopy should prevent death — false, and false in a way no amount of further observational data would have exposed. The error is the assumption that a marker on the causal pathway is a lever on it: ectopy was a symptom of a diseased myocardium, and the drugs suppressed the symptom while making the myocardium worse. This is the surrogate fallacy at maximum strength, with the additional detail that the trial design itself enrolled only responders, which should have made the result more favourable and did not.
My take: I keep this one close because the reasoning is not stupid. It is the reasoning I use. Restore the anatomy and function follows (entry 27), achieve radiographic union and the limb works (entry 50), reattach the tendon and the shoulder recovers (entry 39) — these are the same argument in a different tissue, and cardiology ran the experiment we have mostly not run. The comfort is that this happened in a field with better trials than ours. The discomfort is the same fact.
N Engl J Med 1991;324(12):781–788 · PMID 1900101 · doi:10.1056/NEJM199103213241201 · preliminary report: N Engl J Med 1989;321(6):406–412, PMID 2473403, doi:10.1056/NEJM198908103210629
24

AAOS and ACCP Guidelines for Venous Thromboembolism Prevention Differ: What Are the Implications?

Chest · 2009 · Eikelboom JW, Karthikeyan G, Fagel N, Hirsh J · Comparative guideline critique

What it does: Documents that two expert bodies issued directly conflicting recommendations for thromboprophylaxis after hip and knee arthroplasty from the same trial evidence. The American College of Chest Physicians accepted venographic deep vein thrombosis — largely asymptomatic clot detected on imaging — as a valid efficacy outcome. The AAOS rejected it as an unproven proxy for pulmonary embolism.

Epistemological angle: The clearest demonstration available that evidence underdetermines practice. Neither body was incompetent, neither had different data, and their advice to surgeons was opposite. The entire disagreement reduces to whether one reference class — asymptomatic venographic clot — licenses inference to another, fatal embolism. That is a question about surrogate validity, not about the trials, and no additional trial of the same design could have settled it.
My take: When guidelines conflict, the instinct is to ask which committee was captured or careless. Usually neither. They disagreed about what would count as an answer, which is a deeper disagreement than any about the data and one that more data does not touch.
Chest 2009;135(2):513–520 · PMID 19201714 · doi:10.1378/chest.08-2655
25

Hemiarthroplasty for the Rotator Cuff-Deficient Shoulder

JBJS Am · 2008 · Goldberg SS, Bell JE, Kim HJ, Bak SF, Levine WN, Bigliani LU · Retrospective case series, 34 shoulders

What it does: Reports that 26 of 34 shoulders, or 76%, "satisfied the limited goals criteria described by Neer et al." Forward elevation improved from 78 to 111 degrees, external rotation from 15 to 38. The success threshold was supplied by the surgical tradition that developed the operation.

Epistemological angle: A threshold calibrated to what a procedure can achieve cannot then be used as evidence that it achieves anything. The criterion is not literally unfailable — eight of the thirty-four shoulders did fail it, so the study could have come out worse than it did. What it lacks is any anchor outside the tradition being assessed: nothing but that tradition determines where the bar sits, so a poor result and a good one are both measured against a standard the operation's own advocates set. And the pattern is not isolated: Williams and Rockwood reported 86% satisfactory by this standard in 1996, Sanchez-Sotelo and Cofield 67% in 2001. Three decades of apparently comparable evidence, all resting on a benchmark with no external referent.
My take: Normal function genuinely is the wrong benchmark for a cuff-deficient shoulder, and the alternative genuinely should come from someone who understands the pathology. The loop just closed, and the field cited it to itself for thirty years without noticing that no patient had ever been asked what threshold mattered to them.
J Bone Joint Surg Am 2008;90(3):554–559 · PMID 18310705 · doi:10.2106/JBJS.F.01029 · concept originates with Neer CS 2nd, Craig EV, Fukuda H, JBJS Am 1983;65(9):1232–1244, PMID 6654936
26

Assessment of Spine Surgery Outcomes: Inconsistency of Change Amongst Outcome Measurements

Spine J · 2010 · Copay AG, Martin MM, Subach BR, Carreon LY, Glassman SD, Schuler TC, Berven S · Multicentre PROM analysis, n=460

What it does: Applies four validated outcome measures to the same 460 lumbar surgery patients, each judged against its own established minimal clinically important difference. Only 40.5% of patients were classified consistently as improved or not improved across all four.

Epistemological angle: "Clinically meaningful improvement" turns out to be a property of the instrument rather than of the patient. The same person, at the same visit, is a success on one validated measure and a failure on another. Which means the choice of primary outcome in any trial is a covert decision about who will count as helped — a decision made before any data are collected, usually without discussion, and capable of determining the result.
My take: Read this before the next trial you assess on the strength of its primary endpoint. The endpoint was chosen by people who knew roughly what their intervention was likely to move.
Spine J 2010;10(4):291–296 · PMID 20171937 · doi:10.1016/j.spinee.2009.12.027
27

Surgical vs Nonsurgical Treatment of Adults with Displaced Proximal Humeral Fractures (PROFHER)

JAMA · 2015 · Rangan A, Handoll H, Brealey S et al. · Pragmatic multicentre RCT, n=250

What it does: Compares surgical fixation or replacement against sling immobilisation for displaced proximal humeral fractures involving the surgical neck. No significant difference in Oxford Shoulder Score or quality of life at two years, and similar complication rates. The authors state that the results do not support the trend toward increased surgery.

Epistemological angle: Included here rather than among the trials because of what preceded it. Surgical volume for these fractures had been climbing for a decade on the strength of anatomical reasoning — restore the anatomy and function follows — without a trial supporting it. Mechanistic plausibility functioned as sufficient warrant to adopt, and only a large pragmatic trial was accepted as sufficient warrant to reconsider. The two standards are not the same size.
My take: Restore the anatomy, restore the function is one of the most reasonable-sounding sentences in orthopaedics and it is a hypothesis, not a principle. It has been tested a number of times now and its record is mixed at best.
JAMA 2015;313(10):1037–1047 · PMID 25756440 · doi:10.1001/jama.2015.1629

Also worth reading

Fleming TR, DeMets DL. Surrogate end points in clinical trials: are we being misled? Ann Intern Med 1996;125(7):605–613 · PMID 8815760 · doi:10.7326/0003-4819-125-7-199610010-00011. States the full-mediation criterion formally, and assembles the cases entry 23 dramatises. Most surrogates in orthopaedics have never been assessed against this standard and a good many would fail it.