Rob Parisien, MD, FAAOS

Constrained Permissivism

On the Epistemology of Disagreement: Why two competent surgeons can read the same evidence and rationally disagree.

Rob Parisien, MD, FAAOS · Working paper

Introduction

Disagreement is a normal part of how science makes progress. Competing hypotheses and rival readings of the same data are the engine of inquiry, not necessarily a symptom that someone has blundered. We tolerate and even encourage this because we expect evidence, in time, to sort the better view from the worse. The assumption underneath is one of directionality, almost a teleology, with a classic statement in the American pragmatist philosopher C. S. Peirce: truth is the opinion that inquiry is fated to reach in the long run [19]. Investigation aims at one answer, and present disagreement is a way station on the road to it, provisional, a wager on what further study will confirm.

Yet what looks healthy at the level of a field can look like pathology at the level of two people. When two well-informed surgeons examine the same evidence, understand the same patient, and assign meaningfully different probabilities to the same outcome, there is no comfortable appeal to the long run. The decision is for this patient, and it is now. So we fall back on the individual reading: at least one of them must have reasoned badly. The puzzle this paper takes up is not that disagreement happens. It is whether some disagreement can remain rational even after every factual misunderstanding and every reasoning error has been removed.

That default reading runs deeper than habit. It is built into how we measure care. Beginning in the 1970s and widely influential through the 1990s, John Wennberg and colleagues at Dartmouth showed that the rate of many common operations varied sharply from region to region, far more than differences in illness could explain [17]. To make sense of that, they separated warranted from unwarranted variation: variation not explained by patient illness or by patient preference. The framework is careful, and it does grant that some variation is legitimate. In preference-sensitive care, where the right rate depends on what informed patients choose, variation is warranted so long as it tracks those preferences. The framework has a well-developed place for variation that arises from patient preferences. What it lacks is a comparably clear category for persistent divergence between fully informed clinicians who share the same evidence, the same patient, and the same preferences and still reach different defensible judgments. Variation traceable to the surgeon is unwarranted by default. The assumption that the evidence fixes one right answer is never argued for. It is built in, at the exact level where our question lives: the surgeon's own judgment.

The gap is not peculiar to variation research; it runs through the surgical literature more broadly. When competent clinicians disagree, we usually trace the disagreement to one of three sources. It may reflect error, whether in knowledge, skill, or reasoning. It may reflect missing information that better evidence should eventually resolve. Or it may reflect a difference in what the patient wants. Each source has its own remedy: error should be corrected, missing information gathered, and preference-driven variation respected. This paper identifies a fourth source, one the surgical literature has had little vocabulary to name. Two competent surgeons can share the same evidence, the same patient, and the same patient preferences, make no error, and still reach different defensible judgments about which evidence is most relevant to this patient. Call this genuine evidential underdetermination.

There is also a more practical reason the question cannot wait. The usual reassurance is that even when two clinicians start from different priors (their starting estimates), enough shared evidence will eventually wash the difference out. The merging-of-opinions theorems make this precise: agents who begin from different priors converge, provided they keep updating on the same, endlessly growing body of evidence [18]. Surgery offers neither the endless evidence nor the long run those theorems require. The patient is a single case, the decision is now, and the operation cannot be undone. So the prior never washes out, and the decision is made inside the disagreement rather than after it has resolved.

The aim of the paper, then, is not merely to defend permissivism as an abstract epistemic position. It is to identify a form of clinical disagreement that the surgical literature has largely lacked the vocabulary to describe: disagreement that persists after error, bias, and differences in patient preference have been removed. Recognizing that category matters because it changes how a disagreement should be handled. Disagreement that reflects error should be corrected; disagreement that reflects underdetermination should instead be examined, articulated, and managed. The reference-class argument that follows is the engine, but the reclassification of surgical disagreement is what gives the result its practical point.

In this paper we argue that some disagreement is not error at all. Clinical evidence does not single out one rational prior, because placing a single patient into a reference class requires a judgment that is defensible but not unique. Surgeons already reason with this mix of formal and tacit inference in practice [3]; what follows explains why the resulting differences can be rational rather than mistaken.

The debate: uniqueness vs. permissivism

The persistent disagreement just described is the clinical face of a standing dispute about evidence itself: does a fixed body of evidence permit exactly one rational credence, or a range? Two camps answer differently (Figure 1).

Figure 1. The current debate. Uniqueness: a fixed body of evidence fixes one rational credence, so competent peers must agree. Permissivism: the same evidence permits a range of rational credences, so competent peers may rationally differ.

Uniqueness

Belief aims at truth, and evidence is our guide to it, so the rational thing is to read the evidence’s bearing correctly, no more and no less. On this view a body of clinical evidence about a fracture and a patient has a definite weight and direction. It favors fixation, or healing, by some determinate amount, much as a column of numbers has one correct sum. If two equally informed surgeons, looking at the same images, literature, and patient, reach materially different probabilities of union, then at least one has misread the data. They cannot both be tracking the truth and still disagree. Put sharply: if the same evidence made both 0.8 and 0.4 fully rational, then “rational” would stop meaning “aligned with reality” and start meaning only “internally tidy.” Philosophers press the point with a thought experiment, the arbitrariness objection associated with White [9], and even surgeons who find such devices artificial can usually feel its pull. Imagine a pill that, taken safely, nudges your belief from a credence of 0.7 down to 0.4 while changing nothing about the evidence. If you were truly free to sit anywhere in a permissible band, you should have no real objection to swallowing it, since you would be no less rational afterward. Yet that feels wrong. We do not experience our beliefs as adjustable by mood or chemistry within some tolerated zone. We feel the data should force our hand, and uniqueness takes that feeling at face value.

Permissivism

The trouble is that evidence cannot interpret itself. A lab value, a radiograph, or a trial result carries no weight until it is read through a judgment about which prior cases this patient resembles, and the literature never tells you, uniquely, which patients count as “like this one.” That reference-class judgment partly determines the prior, and there is no provably correct prior to start from. Every attempt to legislate one either generates paradoxes, since the principle of indifference gives different answers depending on how the possibilities are carved, or smuggles in unproven assumptions about how the world is built. So the arbitrariness charge boomerangs. Insisting on one mandatory starting framework is itself the arbitrary move, while allowing a bounded range of defensible ones is the honest response to the fact that reason alone does not hand us our priors. This is not “believe what you like.” A prior must answer to evidence, experience, and mechanism, and most possible priors are plainly out of bounds. And the pill loses its bite. You refuse it not on a whim but because, read through your own defensible framework, 0.4 misrepresents this patient, even as you grant that a colleague who weighs “patients like this” slightly differently could rationally sit at 0.4. The permission here is interpersonal, not a license to move your own credence around at will: a rational credence rests on a stable, reason-responsive framework, and a chemically induced shift that answers to no change in your reasons is not made rational merely because another clinician, reasoning from a different framework, could defensibly hold it. That is the real claim: not one surgeon free to roam a permissible zone, but two surgeons with different defensible priors reaching different rational credences from the same evidence.

The constructive contribution

Start with the reference-class problem. To say how likely an outcome is for a particular patient, you have to treat that patient as a member of some group and borrow that group's rate. But every patient belongs to many groups at once. A 60-year-old with this tibial fracture is also a smoker, also a diabetic, also someone injured in a high-energy crash, and each of those groupings carries a different rate of healing. Which group should this patient's probability be read from? The evidence itself does not say. That is the reference-class problem [1, 5], and it is long familiar: a probability for an individual always rests on a choice of comparison class that the data underdetermine.

The second strand is just as familiar: the active debate, in the epistemology of disagreement, over whether one body of evidence fixes a single rational belief or permits several [2]. We set out those two camps earlier.

Neither idea is new, and to notice either is not the contribution. The contribution is a derivation, not merely a connection: to derive the second from the first. Applied to a single patient, the reference-class problem does not merely allow a range of rational starting beliefs; it forces one.

The forcing is specific to the single case. There is no repeatable frequency for an individual patient, so some judgment about which cases count as relevant is unavoidable. It is not an optional modeling step that a more careful analysis could remove. The argument runs in three steps. First, any clinical probability requires carrying general evidence over to one individual. Second, that carry-over rests on contestable judgments about relevance, similarity, and how well the source evidence transfers. Third, because those judgments are constrained but not settled uniquely, the evidence permits a bounded range of rational priors rather than a single one.

This last step deserves a word of defense, since a determined uniqueness theorist can grant everything so far and still insist that one reference class is the rationally correct one, even if no method can identify it. But a uniquely correct answer that no one’s reasons can reach does no normative work. If two relevance judgments are equally responsive to the available evidence, mechanism, and clinical experience, and no further reason ranks one above the other, then to call only one of them rational is to draw a distinction the agent’s reasons do not support. The plurality of defensible reference classes is therefore a plurality of rationally permissible priors.

The result is constructive, not skeptical. The reference-class problem is often treated as a threat to the very idea of clinical probability. Here it does the opposite work: it explains how rationality can hold a prior to real standards without pinning it to one exact value.

Why relevance has no uniquely determined answer

Relevance has no single right answer for two basic reasons.

The first is a trade-off between fit and sample size. Broad populations give more data but resemble the patient less; narrow subgroups resemble the patient more but are smaller, noisier, and more prone to selection effects. More similarity usually means less data, and less similarity more, and no rule says where the best balance lies.

The second is that the same patient can be described in many equally reasonable ways: by fracture morphology, bone quality, functional demand, injury mechanism, care setting, surgeon experience, or psychosocial context. Each description points to a different comparison group, and evidence can rule some out as unreasonable without crowning one as correct. This is also why "my patient is not the trial population" is so often a legitimate worry rather than a dodge. It is a claim that the trial's group is not the right description of this patient, a question about whether the result transfers, not a refusal to follow evidence [6].

Underlying both is the singularity of the case. We never observe how this patient would have fared under the treatment not chosen, so the estimate has to be imported from other cases, and nothing about the individual settles which cases should dominate. What remains is not unlimited freedom but a bounded range of defensible relevance judgments, and with them a bounded range of defensible priors.

Four responses to the underdetermination

How tightly does rationality constrain the starting prior? Four broad responses are possible, spanning the range from "exactly one" to "anything goes." (A separate question, whether to represent the leftover uncertainty as a single imprecise belief rather than a set of sharp ones, runs orthogonal to this axis; we set it aside here.) The reference-class problem places real pressure on three of them and motivates the fourth, constrained permissivism (Figure 2).

The first answer is that rationality fixes one prior and you can derive it from neutral principles. This is objective Bayesianism: choose the prior that assumes the least, using rules like indifference, symmetry, or maximum entropy [8]. The appeal is obvious, but the approach breaks on a problem that is easier to feel than to argue away. The principle of indifference says that if you have no reason to prefer one value over another, you should spread your belief evenly across them. The catch is: evenly across what? Say you want a neutral starting belief about whether a fracture will unite. You could spread your belief evenly across the probability of union, from 0 to 1. Or you could spread it evenly across the odds of union, the same chance written a different way, as when you say a fracture is four times as likely to heal as not. Both are equally honest ways to say "I have no information," yet they are different priors, and they disagree about how likely union is. A belief that is flat across the probability becomes tilted and informative the moment you rewrite it as odds, and the reverse holds too. So "no information" does not pick out one prior. It does so only after you have chosen which scale to be neutral on, and nothing in the evidence makes that choice for you. Philosophers call this Bertrand-style non-invariance: the supposedly neutral prior shifts when you redescribe the same quantity in an equally legitimate way [7]. The method meant to remove the arbitrary choice of prior smuggles the same choice back in, one level down.

Figure 2. Four responses to how tightly rationality constrains the prior. The reference-class problem places pressure on three of them and motivates constrained permissivism.

Constrained permissivism stated precisely

The view can now be stated exactly. For some bodies of clinical evidence, more than one prior credence is rationally permissible. Facing the same evidence, competent clinicians can rationally settle on different degrees of belief.

The permissible set is not open-ended. Two requirements bound it. The first is coherence: beliefs must update consistently, avoid contradiction, and respond to new evidence in proportion to its evidential force [4]. The second is defensibility: a prior must be grounded in relevant evidence, experience, or mechanism, and able to hold up against the plausible alternatives.

The set has vague edges. Some priors are clearly defensible, some clearly not, and some sit in a contested middle. A boundary can be real without being a sharp numerical line (Figure 3).

Interpersonal permissivism follows directly. Two surgeons can occupy different points in the permissible set, and hold substantially different credences, without either being irrational or factually mistaken. This is not a plea for mutual tolerance. It falls out of the structure of clinical inference itself.

Figure 3. Defensibility narrows the space of priors to a bounded, soft-edged range, not to a single point (uniqueness) and not to nothing (radical permissivism).

Confidence without condemnation

Permissivism is sometimes heard as a call for humility, as if holding that more than one view can be rational means holding any view weakly. It does not. A surgeon can have strong reasons to occupy one part of the permissible range and can hold her prior firmly. What permissivism blocks is not confidence. It is one specific step beyond confidence: the move from "I am highly confident" to "therefore any equally informed colleague who disagrees must be wrong" [9]. Confidence is a fact about the strength of your own position. Uniqueness quietly converts it into a verdict on everyone else's. Permissivism cuts that conversion. So dogmatism, on this view, is not being sure. It is treating your own sureness as proof that every dissenter has reasoned badly.

Objections

One objection grants the reference-class uncertainty but proposes to absorb it into a single representation. Rather than admit a plurality of priors, why not capture the uncertainty in one imprecise object: an interval of probabilities, a model average, a hierarchical model, a sensitivity analysis, or simply suspended judgment [10, 11]? These are all legitimate tools, but none of them removes the plurality. Each still requires judgments about which classes or models to include, how wide to draw the admissible set, how to weight its members, and what counts as relevant. The permissive judgment is relocated, not abolished. And constrained permissivism need not insist that a single precise prior is always best. It claims only that rationality does not require every clinician to represent uncertainty in one mandated form. A working prior, an interval of priors, and a model-averaged estimate can each be permissible.

A second objection comes from the epistemology of disagreement. A competent colleague who has seen the same evidence and landed on a different credence needs to be accounted for. Her disagreement is itself a piece of evidence [13, 16], and it can rightly prompt you to re-examine your assumptions, widen a sensitivity analysis, reconsider what you treated as relevant, or lower your confidence a notch. What it does not do is force both of you to a midpoint. If the evidence genuinely permits more than one relevance judgment, then finding a peer at another permissible point does not show that your own is defective. Constrained permissivism blocks compulsory conciliation, not all of it [14, 15]. The pattern is general: uniqueness tends to support conciliationism, and permissivism tends to support steadfastness. If the evidence fixed a single rational credence, a dissenting peer would have erred, and you would have no special reason to think the error was hers rather than yours, so you should move toward her. But where more than one credence is permissible, a peer at another permissible point need not have erred at all. The steadfastness this licenses is bounded: you may hold a defensible point, but a credence outside the permissible range earns no such protection.

Why are you entitled to hold your point rather than defer? Because the disagreement traces to a permissible difference in epistemic standards. Each clinician may rationally assess the disagreement by the lights of her own defensible standard, the one under which her own point comes out best supported [12]. That is what answers the equal-weight objection, the demand that peers always split the difference. This entitlement is not a neutral vindication of her own view; the standard itself stays open to criticism, to correction by outcomes, and to comparison with the alternatives. Two clarifications keep this from sliding into relativism. First, the claim is local and epistemic, not global and alethic (alethic meaning about truth itself). Rational standards may permissibly differ even though there is a single fact about whether the fracture heals. The view says nothing about that fact, so it needs no "view from nowhere." Second, the standards answer to the same defensibility conditions as the priors. They are constrained, not arbitrary. Divergence within bounds is not relativism.

There is also a tu quoque (a "you too" reply) waiting for anyone who presses the objection hard. To insist that you justify favoring your own standard from some neutral standpoint is to demand the very thing that sank the search for one objectively correct prior. It presupposes a view from nowhere that the rest of the paper has argued does not exist. The challenge can be raised only on an assumption the objector has already given up.

Scope and limits

A few limits keep the claim from being mistaken for a larger one. The argument is about knowledge, not metaphysics. It does not say that there are several physical truths about the patient, that every interpretation is as good as any other, that evidence cannot sharpen judgment, that clinical probabilities are mere personal taste, or that disagreement is always rational. It says only this: the available evidence can permit a bounded range of rational credences, because applying general evidence to a single patient requires relevance judgments that are constrained but not uniquely settled. Whether one outcome is in fact the true one is a separate matter from whether the evidence picks out one rational prior before that outcome is known.

Conclusion

Constrained permissivism is, in the end, a claim about what clinical disagreement can be. The surgical literature has serviceable accounts of disagreement that comes from error, from missing information, and from differing patient preferences. It has lacked an account of disagreement that comes from genuine evidential underdetermination, where competent surgeons reason well, share the same evidence and the same patient, and still differ. Supplying the conceptual basis for that fourth category is the practical contribution of this paper.

It matters because the categories call for different responses. Surgical disagreement is not a single phenomenon. Some of it reflects error and should be corrected; some reflects underdetermination and should instead be examined, articulated, and managed. Once that distinction is in view, it reframes several familiar settings. In peer review, disagreement with a reviewer does not by itself establish substandard reasoning. In consultation, the aim may be to surface a colleague’s different reference-class judgment rather than to find the surgeon who erred. In a multidisciplinary conference, a disagreement becomes more productive when each clinician is asked which patients, mechanisms, or outcomes they are treating as most relevant. In shared decision-making, a patient can be told honestly that competent surgeons may continue to disagree, rather than offered the fiction that one expert must hold the answer. And in variation research, divergence traceable to clinician judgment should not be filed under unwarranted by default.

Working any of these out in full is a separate task, and not one this paper completes. What it offers is the thing they all depend on: a principled reason to hold that some disagreement among competent surgeons is neither error to be corrected nor noise to be averaged away, but a rational consequence of how general evidence meets a single patient.

References

[1] Hájek, A. (2007). The reference class problem is your problem too. Synthese, 156(3), 563–585.

[2] Kopec, M., & Titelbaum, M. G. (2016). The uniqueness thesis. Philosophy Compass, 11(4), 189–200.

[3] Parisien, R., Drost, A., Razi, A., Ramtin, S., Ring, D., & Janssen, S. J. (2026). Musculoskeletal surgeons use mixed reasoning rather than pure Bayesian strategies in clinical practice. PLOS ONE, 21(6), e0351694.

[4] de Finetti, B. (1974). Theory of Probability. New York: Wiley.

[5] Reichenbach, H. (1949). The Theory of Probability. Berkeley: University of California Press.

[6] Bareinboim, E., & Pearl, J. (2016). Causal inference and the data-fusion problem. Proceedings of the National Academy of Sciences, 113(27), 7345–7352.

[7] van Fraassen, B. C. (1989). Laws and Symmetry. Oxford: Oxford University Press.

[8] Jaynes, E. T. (2003). Probability Theory: The Logic of Science. Cambridge: Cambridge University Press.

[9] White, R. (2005). Epistemic permissiveness. Philosophical Perspectives, 19, 445–459.

[10] Walley, P. (1991). Statistical Reasoning with Imprecise Probabilities. London: Chapman & Hall.

[11] Joyce, J. M. (2005). How probabilities reflect evidence. Philosophical Perspectives, 19, 153–178.

[12] Schoenfield, M. (2014). Permission to believe: Why permissivism is true and what it tells us about irrelevant influences on belief. Noûs, 48(2), 193–218.

[13] Christensen, D. (2007). Epistemology of disagreement: The good news. Philosophical Review, 116(2), 187–217.

[14] Elga, A. (2007). Reflection and disagreement. Noûs, 41(3), 478–502.

[15] Kelly, T. (2010). Peer disagreement and higher-order evidence. In R. Feldman & T. A. Warfield (Eds.), Disagreement (pp. 111–174). Oxford: Oxford University Press.

[16] Christensen, D. (2010). Higher-order evidence. Philosophy and Phenomenological Research, 81(1), 185–215.

[17] Wennberg, J. E. (2010). Tracking Medicine: A Researcher’s Quest to Understand Health Care. New York: Oxford University Press.

[18] Blackwell, D., & Dubins, L. (1962). Merging of opinions with increasing information. Annals of Mathematical Statistics, 33(3), 882–886.

[19] Peirce, C. S. (1878). How to make our ideas clear. Popular Science Monthly, 12, 286–302.