Rob Parisien, MD, FAAOS

Two Readings of “Surgery Causes Union”

Causal Ambiguity in Constructed Surgical Endpoints

Rob Parisien, MD, FAAOS · Working paper

1. An ordinary claim

Surgeons say that the plate caused the fracture to unite. They say that fixation caused union by twelve weeks, or that a construct failed to cause healing. The claims pass without comment. They appear in Methods, are tabulated in Results, and are pooled across trials in meta-analyses as though the thing being asserted were stable from one study to the next. On the surface they are the least controversial statements in the field: a physical intervention produced a physical-biological outcome. What could be less ambiguous than a bone that knits, or fails to?

Ask, though, what the surgery is a cause of. Two answers are available, and ordinary surgical language does not choose between them. On the first, surgery causes a process in the tissue — the cascade of biological events that ends in a healed bone. On the second, surgery causes a judgment — the clinician’s registration, on a radiograph or a scoring form, that union has occurred. Nothing in the sentence “surgery caused union” tells us which claim is being made, and in practice the two are used interchangeably, often within a single paper.

This essay tries to make that ambiguity visible and to show that it is not merely loose talk. The distinction between the two readings has consequences at specific, recognizable points in how outcomes are reported, how thresholds are chosen, and how trials are synthesized. Where the endpoint is constructed — where what counts as the outcome depends on a scoring system or an interpretive convention — the ambiguity is not decorative. It is doing classificatory work that the clinician does not see themselves doing.

The claim here is deliberately narrow. It is not the familiar thesis that medicine needs a pluralism of causal concepts in general — Cartwright’s (2007) point that causal notions are diverse and resist reduction to a single relation, or the Russo–Williamson (2007) insistence that credible causal claims in the health sciences rest on both mechanistic and difference-making evidence, a demand now institutionalized in the EBM+ program (Clarke et al. 2014). Those debates form the backdrop. The argument advanced here is more local: a specific class of surgical endpoints carries a specific ambiguity, one that has gone largely unnoticed precisely because the endpoints look so concrete.

2. Two readings

Consider the two readings side by side, because they are frequently mistaken for two ways of measuring one thing when they are in fact two different claims.

The process reading. On this reading, “surgery causes union” picks out a physical process: load-sharing across the fracture site, the periosteal and endosteal response, vascular ingrowth, the organization of callus, and eventual remodeling. Its truth conditions live at the level of tissue. Whether they are satisfied does not depend on anyone looking; the bone heals or it does not. The natural theoretical home for this reading is the process tradition in the theory of causation — Salmon’s (1984) account of causal processes and Dowe’s (2000) refinement of it as the transmission of conserved quantities through continuous spatiotemporal processes. The referent is mind-independent in the strong sense: a fact about matter, indifferent to observation.

The judgment reading. On this reading, “surgery causes union” picks out an epistemic act: a radiograph read as showing cortical continuity, a scoring threshold crossed, a line written in the chart. Its truth conditions live at the level of the clinician’s registration of a state. Its natural home is counterfactual. In Lewis’s possible-worlds analysis (1973a, 1973b), surgery causes union just in case, in the nearest world where the surgery does not occur, the clinician does not judge union. The same counterfactual can be regimented in Pearl’s (2009) structural framework, where it is evaluated by intervening on the surgical variable and reading off the effect on the recorded outcome — which is, as it happens, the idiom that trial-based inference implicitly runs on. Either way, the relatum is not the bone as such but the assessment of the bone.

There is a parallel here worth naming, because it shows the ambiguity is not a quirk of surgery. Hall (2004) argues that our single word “cause” answers to two distinct concepts: production, a local, physical notion of one event generating another, and dependence, the counterfactual notion that one event would not have occurred without the other. The two readings line up with Hall’s two concepts — the process reading is causation-as-production, the judgment reading is causation-as-dependence. The alignment is not exact: Hall’s split concerns two concepts of a relation between the same events, whereas the split here concerns two different relata, tissue on one side and assessment on the other. But the parallel is instructive. The capacity of one causal verb to carry two claims is a general feature of causal language, not a local defect; “cause union” is simply a place where that feature acquires clinical teeth.

These are not two instruments trained on the same target. The process at the tissue level can occur without anyone registering it — the silent unions found incidentally on imaging obtained for other reasons. And the judgment can occur without the full process — a threshold crossed on a scale before remodeling is complete, an assessment made prematurely but not incorrectly by the standards of the scale. The relationship between process and judgment is empirical and imperfect: a matter of how well the assessment tracks the tissue, which can be tight in one setting and loose in another. It is not a definitional identity, and treating it as one is the mistake that generates most of what follows.

Anscombe’s (1971) observation about causal verbs belongs here. Verbs like push, carry, and scrape each name a family of distinct physical relations rather than one; “cause” inherits this variety, and in constructions like “cause union” it smuggles more than one causal notion into a single word. What looks like one predicate is doing at least two jobs.

3. Why the collapse persists

If the two readings are genuinely distinct, why does the collapse between them feel so natural — natural enough that pointing it out can seem pedantic? Three considerations explain its durability, and the third explains why it is not simply an error to be corrected.

Folk essentialism about union. Surgeons, journals, and guidelines tend to treat union as a natural kind with determinate boundaries — a real thing whose edges are fixed independently of us, which scoring systems merely detect. Under this assumption the two readings collapse trivially, because the judgment is just the process plus measurement noise: get the imaging good enough and the gap closes to zero. This is a form of the psychological essentialism that cognitive and developmental psychologists have documented as a default mode of categorization — the tendency to posit hidden essences behind observed features (Gelman 2003). Mayr (1982) called the historical version the “dead hand of Plato”: the typological habit of treating variation as deviation from an underlying type rather than as the reality itself. When union is imagined this way, the question “process or judgment?” barely registers, because the judgment is assumed to be a transparent window onto the process.

Interactive kinds. But union is not a natural kind of that sort. It is closer to what Hacking (1995, 1999) called an interactive kind, subject to looping effects. The scoring systems — RUST, the modified RUST, various radiographic union criteria — are not passive instruments. They shape what counts as union; that shapes how residents are trained to read radiographs; that shapes which trials get designed and how their endpoints are operationalized; and that feeds back into the next revision of the scoring systems. The category and its instruments co-evolve. Boyd’s (1999) homeostatic property cluster account gives the more accurate picture of what union is: the radiographic, mechanical, histological, and clinical features that tend to co-occur in a healing bone form a cluster held together by underlying mechanisms, not a kind with sharp edges. Treating that cluster as though it had crisp boundaries is precisely what manufactures the appearance of a single stable referent — and with it, the appearance that the two readings must coincide.

The collapse as a productive fiction. None of this makes the collapse stupid. It is useful. Treating the judgment as a transparent readout of the process is what lets surgeons communicate quickly, trials define endpoints at all, meta-analyses pool results, and clinics run. A working category with usable edges is worth a great deal, and the demand that every clinician hold the distinction in mind at every moment would be both impractical and pointless. The argument is not that surgeons should stop talking this way. It is that the collapse conceals a real ambiguity, and that the concealment stops being harmless at particular pressure points — places where the difference between the readings becomes consequential.

4. Where it bites

Discordance between Methods and Results. Methods sections operationalize the judgment reading. They specify the scoring system, the threshold, the imaging protocol, the timing of assessment — all the machinery of registering a state. Results sections, and still more Discussion sections, frequently slide into the language of process: the construct “promoted healing,” the fixation “accelerated union,” the technique “improved bone formation.” Anyone who reads the trauma literature closely will recognize the pattern, and it is easy to dismiss as careless writing. It is not merely that. The drift is the surface symptom of treating two distinct causal claims as one. A paper that measured a judgment in its Methods and reports a process in its Discussion has changed the referent of its central claim without announcing it, and usually without its authors noticing.

Thresholds as commitments about which reading you are making. A scoring threshold looks like a methodological detail — a number chosen for convenience or convention. It is more than that. A threshold set at, say, eleven on one radiographic scale, fourteen on another, and ten on a third is not three ways of drawing the same line. Each choice carries an implicit commitment about how tightly the judgment tracks the process. A more permissive threshold makes union easier to register but widens the gap between the registration and the underlying tissue state; a more conservative threshold tightens the coupling but pays for it in false negatives — bones that have healed but not yet been credited. To choose a threshold is therefore to take a stance on the causal claim being made: on how much of the process one is willing to read off the judgment. The choice is not upstream of the causal question. It is the causal question, wearing methodological clothing.

The same structure in arthroplasty. The pattern is not confined to fracture care. In joint arthroplasty, the minimal clinically important difference (MCID) has the same architecture, and a recent commentary in the arthroplasty literature identifies the problem without naming it as a causal-semantic one. “Surgery caused a meaningful improvement” divides exactly as “surgery caused union” does: did the surgery cause a functional change in the patient’s body and life, or did it cause the patient or clinician to register a change crossing a defined threshold on an instrument? The MCID literature is plainly reaching for this distinction — it has the vocabulary of “meaningful,” “detectable,” “patient-acceptable” — but that vocabulary cannot quite hold the difference between a change in the person and a change in the score. The construct differs; the ambiguity is identical.

Making the dependencies explicit. There is a positive move implied by all of this. Consider what changes when an analyst stops treating the recorded endpoint as a transparent readout of the biological process and instead makes the interpretive commitments inside the endpoint visible — for instance, by conditioning the analysis on stated assumptions about how tightly judgment tracks process, and reporting how the conclusion moves as those assumptions vary. Pearl’s (2009) framework is one way to write those dependencies down explicitly; Bayesian priors are another; the point depends on neither. What matters is the move: refusing to let the judgment reading pass as a transparent measurement of the process reading, and putting the coupling on the table where it can be argued about.

5. Scope

A caveat is owed, because the argument would overreach if left unbounded. It is not the claim that all causal statements in surgery are ambiguous in this way. Many are not. Mortality, reoperation, deep infection requiring debridement, hardware failure visible on imaging — these have mostly stable process-level referents, well coupled to their judgment-level assessments. When a paper says surgery caused a death or a reoperation, the two readings all but coincide, and nothing is hidden by running them together.

The ambiguity bites where the endpoint is constructed — where what counts as the outcome depends on a scoring system, a threshold, or an interpretive convention that is doing classificatory work invisibly. Union is the central case, but it is a member of a class, and MCID is another member. The essay is therefore not a wholesale indictment of surgical causal inference. It is an analysis of a specific failure mode that afflicts a specific kind of endpoint, with concrete implications for how such endpoints are reported, how trials about them are designed, and how guidelines synthesize them.

That yields a single, practicable demand. When a study claims that surgery caused a constructed endpoint, the reader is owed enough information to evaluate both readings independently. The Methods should specify the judgment: the instrument, the threshold, the protocol. The Discussion should not silently migrate to the process. And where a study genuinely wants to make a process-level claim — that the intervention changed the biology of healing, not merely its registration — it should defend the coupling between judgment and process rather than assume it. The distinction, once seen, is not hard to state. The difficulty is entirely in noticing that there was a distinction to honor at all.

References

Anscombe, G. E. M. (1971). Causality and Determination. Cambridge: Cambridge University Press.

Boyd, R. (1999). Homeostasis, species, and higher taxa. In R. A. Wilson (Ed.), Species: New Interdisciplinary Essays (pp. 141–185). Cambridge, MA: MIT Press.

Cartwright, N. (2007). Hunting Causes and Using Them: Approaches in Philosophy and Economics. Cambridge: Cambridge University Press.

Clarke, B., Gillies, D., Illari, P., Russo, F., & Williamson, J. (2014). Mechanisms and the evidence hierarchy. Topoi, 33(2), 339–360.

Dowe, P. (2000). Physical Causation. Cambridge: Cambridge University Press.

Gelman, S. A. (2003). The Essential Child: Origins of Essentialism in Everyday Thought. Oxford: Oxford University Press.

Hacking, I. (1995). The looping effects of human kinds. In D. Sperber, D. Premack, & A. J. Premack (Eds.), Causal Cognition: A Multidisciplinary Debate (pp. 351–394). Oxford: Clarendon Press.

Hacking, I. (1999). The Social Construction of What? Cambridge, MA: Harvard University Press.

Hall, N. (2004). Two concepts of causation. In J. Collins, N. Hall, & L. A. Paul (Eds.), Causation and Counterfactuals (pp. 225–276). Cambridge, MA: MIT Press.

Lewis, D. (1973a). Causation. Journal of Philosophy, 70(17), 556–567.

Lewis, D. (1973b). Counterfactuals. Cambridge, MA: Harvard University Press.

Mayr, E. (1982). The Growth of Biological Thought: Diversity, Evolution, and Inheritance. Cambridge, MA: Harvard University Press.

Pearl, J. (2009). Causality: Models, Reasoning, and Inference (2nd ed.). Cambridge: Cambridge University Press.

Russo, F., & Williamson, J. (2007). Interpreting causality in the health sciences. International Studies in the Philosophy of Science, 21(2), 157–170.

Salmon, W. C. (1984). Scientific Explanation and the Causal Structure of the World. Princeton: Princeton University Press.

[Arthroplasty MCID commentary referenced in §4 — recent Journal of Bone and Joint Surgery commentary; full citation to be supplied by author.]