Theme Six
What observation can and cannot buy
What does allocation buy you, and what can a cohort do without it?
Randomisation is often treated as a property a trial either has or lacks. It is better understood as a loan against future adherence: when patients and surgeons cross between arms in large numbers, the loan is called in. But the converse lesson is the one this theme exists to correct. Observational evidence is not a weaker version of a trial — it answers questions no trial can reach, and where it fails it usually fails for a specific and correctable reason rather than because it lacked randomisation.
28
SPORT: The Trial and Its Observational Twin
JAMA · 2006 · Weinstein JN, Tosteson TD, Lurie JD et al. · RCT (n=501) and parallel observational cohort (n=743), published the same day
What it does: Randomises patients with lumbar disc herniation to surgery or non-operative care. By three months only half the surgical arm had been operated on and 30% of the non-operative arm had crossed to surgery. The intention-to-treat analysis showed small, non-significant differences, and the authors state that conclusions about superiority or equivalence are not warranted from it. The companion paper follows those who declined randomisation, analysed as treated, and finds a clear significant advantage for surgery on every primary outcome.
Epistemological angle: The best teaching pair in this collection. Same investigators, same outcome measures, same eligibility criteria, same journal, same day — and abandoning randomisation for observed treatment choice produces a dramatically cleaner and more confident-looking answer. Which of the two papers a surgeon cites is a reliable indicator of what they believed before opening the journal. The observational result is not fabricated; it is what confounding by indication looks like when the people choosing surgery are systematically different from those who do not.
My take: A well-funded, well-run, adequately powered trial that cannot answer its own question. That is worth sitting with, because the usual explanations for disappointing trials — too small, too sloppy, wrong outcome — do not apply here. Adherence is a precondition for randomisation to mean anything, and in surgical trials it is frequently unobtainable.
29
Hormone Therapy and Coronary Heart Disease: The Cohort, the Trial, and the Reconciliation
JAMA · 2002 · Rossouw JE, Anderson GL, Prentice RL et al. (Women's Health Initiative), RCT n=16,608 · with Epidemiology · 2008 · Hernán MA, Alonso A, Logan R, Grodstein F, Michels KB, Willett WC, Manson JE, Robins JM · re-analysis of the Nurses' Health Study
What it does: A large, careful, decades-long observational cohort had found that women taking combined hormone therapy had less coronary heart disease. The randomised trial found the opposite: a hazard ratio of 1.29 for coronary events, with stroke and pulmonary embolism also raised, and the trial stopped early because harms exceeded benefits. Six years later Hernán and colleagues re-analysed the original cohort as though it were a sequence of trials, emulating the randomised design and its intention-to-treat analysis, and largely reproduced the trial's estimates from the same observational data.
Epistemological angle: Included as the general-medicine counterpart to entry
28, and it teaches something that pair cannot. The first two papers look like the familiar story that observational evidence is unreliable and randomisation is the remedy. The third undoes that reading. The cohort's data were not wrong; the analysis had compared prevalent users against never-users, which selects for women who had already tolerated the drug, and had aligned time zero incorrectly. Re-analysed with the trial's own design imposed on it, the cohort gave the trial's answer. So the failure was not observational data as a category but a specific and correctable analytic choice — which means the lesson is not "distrust cohorts" but "ask what trial this analysis is emulating, and whether anyone specified it." That is a far more useful question to carry into a registry paper, and it is the question entry
30 passes and most do not.
My take: This is the entry I would put in front of anyone who says the words "real-world evidence" approvingly. The gap between the cohort and the trial was not a story about randomisation's magic. It was a story about a design question nobody had been required to ask out loud, and about how long it took — six years and a good methodologist — to find out that the data had been able to answer correctly all along.
30
Failure Rates of Stemmed Metal-on-Metal Hip Replacements: Analysis of the National Joint Registry of England and Wales
Lancet · 2012 · Smith AJ, Dieppe P, Vernon K, Porter M, Blom AW · Registry cohort, n=402,051 primary hip replacements
What it does: Analyses four hundred thousand hip replacements and finds that stemmed metal-on-metal implants failed at markedly higher rates than alternatives, with risk rising with head size and concentrated in younger women.
Epistemological angle: The positive case for observational data, and it turns entirely on class size. The signal was slow, dose-dependent and modified by patient characteristics — invisible to any individual surgeon's experience and beyond the power of any feasible randomised trial. The effect modification by sex and head size is the part a trial would most likely have missed even if one had existed. Corroboration from the Australian registry on a different population is what converted a worrying pattern into a conclusion. Note the precondition the next two entries remove: the registry is compulsory, national, and held by nobody who sells implants.
My take: Worth holding alongside the sham trials as a corrective. The lesson of this collection is not that observational evidence is weak and trials are strong. It is that each answers questions the other cannot, and knowing which question you are asking is the whole skill.
Lancet 2012;379(9822):1199–1204 · PMID 22417410 ·
doi:10.1016/S0140-6736(12)60353-5 · corroborated by de Steiger RN et al., JBJS Am 2011;93(24):2287–2293, PMID 22258775
Also worth reading
Montori VM, Guyatt GH. Intention-to-treat principle. CMAJ 2001;165(10):1339–1341 · PMID 11760981 · PMC81628 · no DOI in record. Short, free, and the reason entry 28's two papers are one valid comparison and one that has quietly reverted to observational status. Intention-to-treat is not a conservative convention but the condition under which a randomised comparison retains any causal warrant at all.
Katz JN, Brophy RH, Chaisson CE et al. Surgery versus physical therapy for a meniscal tear and osteoarthritis (METEOR). N Engl J Med 2013;368(18):1675–1684 · PMID 23506518 · doi:10.1056/NEJMoa1301408. Crossover designed in rather than suffered: 30% of the physiotherapy arm had surgery within six months, so the null is equally consistent with equivalence and with a third of that arm having received the comparator.