Theme Seven
What survives the retelling
What survives the compression from data to conclusion sentence, and who decided what went in?
Almost nobody reads trials. Surgeons read abstracts, and mostly the last sentence of them. This theme follows one continuum from end to end: outcomes forgotten between protocol and paper, conclusions framed to matter, hedges present in the Discussion and absent from the Abstract, and — at the far end, where the same failure stops looking accidental — data the reader never had the opportunity to see at all. The intensity varies. The structure does not.
31
Empirical Evidence for Selective Reporting of Outcomes in Randomized Trials
JAMA · 2004 · Chan AW, Hróbjartsson A, Haahr MT, Gøtzsche PC, Altman DG · Protocol-to-publication cohort, 102 trials, 3736 outcomes
What it does: Compares ethics-committee-approved protocols against the papers eventually published. Half of efficacy outcomes and 65% of harm outcomes were incompletely reported. Statistically significant outcomes were two to five times more likely to be fully reported. 62% of trials had at least one primary outcome changed, added or dropped. When the trialists were surveyed, 86% denied that unreported outcomes existed.
Epistemological angle: The field-defining paper, and the last figure is the one that matters. The investigators were not concealing results; they had forgotten. Selective reporting is mostly not misconduct but ordinary cognition operating on data the analyst has already seen — outcomes that behaved interestingly feel like the real findings, and the others fade. This is why procedural remedies like pre-registration are necessary: the failure mode is invisible from inside the person committing it.
My take: It seems safest to assume the published literature leans towards benefit and away from harm — not through anyone's dishonesty, but because that is the shape of the filter it passes through.
32
Analyzing Spin in Abstracts of Orthopaedic RCTs with Statistically Insignificant Primary Endpoints
Arthroscopy · 2020 · Arthur W, Zaaza Z, Checketts JX et al. · Meta-research, 250 RCTs screened
What it does: Examines abstracts of orthopaedic trials whose primary endpoint was null. 44.8% contained spin; among those, 79.5% had it in the conclusion. JBJS showed the highest prevalence at 56.8%. There was no association with industry funding.
Epistemological angle: Two findings, and the second is the interesting one. Spin concentrates in the conclusion sentence — the part actually read, quoted and remembered — so the distortion is maximal at the point of highest readership. And the absence of a funding association undercuts the comfortable story that this is a money problem. It is a wanting-your-work-to-matter problem, which is far more widespread and much harder to legislate against.
My take: Note that this sits awkwardly beside a wider meta-research literature in which industry funding does predict favourable conclusions. I have not resolved the tension and would rather leave it standing than pick a side: competent groups asking adjacent questions have got different answers about the role of money.
33
Epistemic Asymmetry Between Abstracts and Discussion Sections in the Orthopaedic Literature
JBJS · 2026 · Parisien R · Quantitative linguistic corpus analysis, 201 JBJS and CORR publications
What it does: Measures the epistemic stance of the same paper in two places. Across 201 orthopaedic publications, Abstracts hedge markedly less than the Discussion sections of the papers they summarise, with a large effect size (Cohen's d = 1.15). A companion measure found that 94% of Abstracts surface no limitation at all. Open-access and paywalled papers did not differ, which locates the asymmetry in the genre rather than in who can read it.
Epistemological angle: Entry
32 shows that some abstracts overstate null results. This asks a prior question of every abstract, null or positive: how much of the paper's own stated uncertainty survives into the part that gets read. The answer is that the qualifications are written and then removed at the point of compression — the authors know what their study cannot support and say so where almost nobody looks. That reframes the problem. It is not that the field fails to acknowledge uncertainty, which entry
36 might be read as showing; it is that acknowledgement is structurally quarantined in the section with the smallest readership. Whether that constitutes a distortion depends on what one thinks an Abstract is for, and the paper does not settle that question.
My take: The 94% is the number worth arguing about, and the obvious objection is that an Abstract has a word limit and is not the place for caveats. That is fair as far as it goes, but the same limit does not prevent a stated effect size, so the omission is a choice about what the space is for. The study measures hedging language rather than whether readers were actually misled, which is the question one would want answered and a harder one to design. What I can say is that writing it changed how I read my own abstracts.
J Bone Joint Surg Am 2026 · in press; volume, pages, PMID and DOI to be supplied at publication · a commissioned Commentary & Perspective accompanies it
34
VIGOR and the Expression of Concern
NEJM · 2000 · Bombardier C, Laine L, Reicin A et al. (VIGOR), RCT n=8,076 · and NEJM · 2005 · Curfman GD, Morrissey S, Drazen JM, Expression of Concern
What it does: VIGOR reported that rofecoxib caused fewer serious gastrointestinal events than naproxen, and noted in passing a higher rate of myocardial infarction in the rofecoxib arm (0.4% against 0.1%), attributed in discussion to a protective effect of naproxen. Five years later the journal's editors published an Expression of Concern stating that additional myocardial infarctions known to at least some authors before publication had not appeared in the paper.
Epistemological angle: The custody problem stated outside orthopaedics, and placed before the two entries that state it inside. The failure here is not that peer review missed a subtle signal; the harm data existed and the reviewing journal did not have them. That locates the vulnerability precisely: peer review evaluates a manuscript, and a manuscript is a selection from a dataset made by a party with an interest in the selection. Note also the asymmetry of framing within the paper itself — the same numerical contrast was read as naproxen protecting rather than rofecoxib harming, which is a live demonstration of how much work the choice of comparator does when a result is turned into a sentence.
My take: I include this partly so that entry
35 is not read as a story about orthopaedics. It is a story about what happens when the party that owns the data also owns the product, and that arrangement is not ours. Ours is simply one of the places it has been documented.
35
The BMP-2 Sequence: What the Published Record Showed and What the Data Showed
Spine J · 2011 · Carragee EJ, Hurwitz EL, Weiner BK, comparison of 13 industry-sponsored trials against regulatory data · and Ann Intern Med · 2013 · Fu R, Selph S, McDonagh M et al., independent individual-patient-data reanalysis
What it does: Carragee and colleagues compared what thirteen industry-sponsored trials of rhBMP-2 reported against what regulatory submissions and later follow-up showed. None of the thirteen published papers reported any adverse event attributable to the product; comparison with other sources suggested true rates ten to fifty times higher. Two years later Fu and colleagues obtained the underlying individual patient data through the Yale Open Data Access project and reanalysed it independently, finding no clinical advantage over iliac crest bone graft and increased harms including cancer risk, with the early publications having misrepresented both.
Epistemological angle: The most disturbing pair here and a necessary one. Peer review functioned, publication proceeded, specialty consensus formed, and the harm signal sat in a regulatory file nobody was reading. "Published in a good journal" is a statement about a process, not a warrant of truth, and the process has a specific blind spot: it evaluates the manuscript in front of it rather than the data behind the manuscript. What eventually settled the dispute was not more argument about the published papers, more editorials, or more expert opinion, all of which had been tried. It was independent access to the raw data. Transparency in this story is not an ethical nicety appended to good science; it was the only mechanism capable of terminating the disagreement.
My take: Thirteen trials and zero adverse events. With hindsight that figure is hard to credit, since no surgical intervention is without complications; at the time it read as a good safety profile, which is how such figures usually read. The implication of the pair is uncomfortable and I have not found a way to soften it: where the sponsor controls the data, the published literature may not be sufficient evidence on its own, however many trials it contains.
Also worth reading
Okike K, Kocher MS, Wei EX, Mehlman CT, Bhandari M. Accuracy of conflict-of-interest disclosures reported by physicians. N Engl J Med 2009;361(15):1466–1474 · PMID 19812403 · doi:10.1056/NEJMsa0807160. 71.2% of payments disclosed overall but only 50.0% of indirectly related ones. A disclosure statement reports which relationships the author judged relevant, which is a different and less useful thing than which existed.
Ruelos VCB, Masood R, Puzzitiello RN et al. The reverse fragility index. Knee Surg Sports Traumatol Arthrosc 2023;31(8):3412–3419 · PMID 37093236 · doi:10.1007/s00167-023-07420-0. Median three outcome events would flip a reported null to significant, and in 81.3% of studies loss to follow-up exceeded that number.