On the epistemic bases of clinical prior credence in surgery
Abstract. Bayesian accounts treat the clinical prior as a single object. I argue that what the formalism flattens into one prior is a compound of distinct epistemic bases—chief among them reference-class and sedimented (the tacit residue of experience)—and that a prior’s basis determines what it takes to answer for it. This overturns three defaults in how evidence is read: that the published result is the verdict, that the surgeon’s expectation is mere bias, and that disagreement between careful readers is error. The result is a constrained subjectivism: a prior is subjective but not arbitrary, because it is answerable.
Consider a common clinical scenario: before an operation, facing the patient in the examination room, a surgeon needs an answer to some version of the question: “Given what I am seeing, how likely is it that this patient heals, or that this operation helps?” The question is about this patient and no one else, and it asks for something less obvious than it sounds: a degree of confidence, as opposed to a fact one could look up. The natural place to look for an answer is the published research, and it generates two kinds of numbers that look like answers: a frequency—i.e. a study reports that 80% of patients improved—and a verdict—the difference reached significance at p < 0.05. Neither is the answer the question asks for.
The 80% looks like a fact about the world—a frequency, the rate at which patients like her improve over a long run of similar cases—but whatever its metaphysical status, it does not settle the question. A group rate is a property of a reference class, and this patient belongs to many, each defining a different rate; moving from any of them to a warranted expectation about her requires choices the frequency does not supply. The number the question wants is not a population rate but a degree of confidence about a single case (Figure 1). I adopt the subjectivist interpretation of probability, on which a probability is a degree of belief rather than a feature of the world [1]. I take subjectivism here as the working framework, not as an established result; what the argument needs is only the weaker claim that a group frequency underdetermines the credence warranted for an individual—which holds on any interpretation, since even a committed objectivist about chance must still choose which chance, which class, applies to this patient.
The question the clinician brings—how likely is the outcome, given what I am seeing?—has a formal and familiar expression, P(H | D), but the numbers most studies report answer different questions. The p-value answers a narrow one: how surprising data this extreme, or more so, would be under the null model—P(D | H0), not the P(H | D) the clinician wants [2]. Turning one into the other requires something the data do not contain, a prior sense of how plausible the effect was to begin with; the same p < 0.05 warrants more confidence for a mechanistically grounded treatment than for an implausible one, and what separates the two is nothing in the trial. Bayes’ theorem is simply the rule that joins them: the post-test belief is the pre-test belief, strengthened or weakened in proportion to how well it accounts for the evidence just observed. The point is not that one school of statistics is right and another wrong, but that a prior is doing work in any reading of evidence that moves from data to a conclusion, whether or not it is named—a point pressed at the scale of whole literatures by Stegenga, whose case for medical nihilism is at bottom an argument about which prior the track record of medical research warrants [29]. My question sits upstream of any particular assignment: not where the prior should be set, but what kind of thing it is. None of this is entirely alien to the surgeon, whose expectation of how a patient will do follows (normatively) a Bayesian arc even if never stated; the starting point of that arc, the prior, is my subject.

Figure 1. The same 80%, read two ways. Read as a population frequency, it is the rate at which patients like her improve over the long run; but she belongs to many such groups at once, and the group rate does not by itself determine the credence warranted for this patient. The number the bedside question wants is an individual credence—a degree of belief located in the reader rather than the population (right). A prior is a probability of this second kind.
The canonical form of Bayes’ theorem leaves something unspecified. It treats the prior as a single, stateable distribution, and is silent about what kind of thing that distribution is. The silence is easy to miss, because the formalism shows two of the prior’s properties plainly—its strength, how far it will let the data move the posterior, and its location—while concealing a third: its basis, where the credence came from, and so what could justify or correct it. I argue that the concealed basis is not one thing but two—reference-class and sedimented—and that a prior’s basis determines what it takes to answer for it. That is the load-bearing claim. Not that priors are subjective, which they are, nor that they differ in strength, which they do, but that they are accountable in different ways and on different timescales according to their basis—and that this is what separates warranted clinical judgment from bias, without either deferring to the formalism or dismissing the expertise that never took the form of a stateable rule.
Sorting priors this way is not bookkeeping, and the difference it tracks is kind, not strength. It overturns three defaults that govern how a trial gets read at the bedside: that the published result is the verdict—its estimate the answer, its significance a yes-or-no; that the surgeon’s own expectation is the bias to be disciplined away; and that disagreement between careful readers is error. These three corrections, not the taxonomy that yields them, are the point. The position that results is a constrained subjectivism: a prior is subjective but not arbitrary, because it is answerable.
Consider two bases on which a clinician’s prior about a given patient might rest, each a credence the clinician actually holds. The first is reference-class: a base rate applied from a representative group or a published cohort—say an 80% rate of meaningful improvement on an outcome score, the very number the opening patient was offered—one anyone can read without ever having examined the patient. The second is sedimented: the clinician’s own sense, on looking at this case, of how this patient will fare—the surgeon who sees a fracture pattern and feels, before they can state their justifications, that this one will not hold, or that a particular elderly patient will sail through surgery others would judge too frail. It is a confidence laid down by a career of having seen, treated, and followed thousands of patients, and is more available than the reasons for it.
A third kind of thing is often called a prior, and naming it marks the boundary of what follows. A rule-mediated prior—a validated risk score, a guideline default, a flat prior chosen by methodological convention [4,5]—names not a separate epistemic basis but a mode of construction. Adopted as a genuine degree of belief, its warrant runs back to whatever the rule encodes—a population, a body of evidence, a mechanism—so it answers as a reference-class prior whose class the procedure has frozen; unadopted, it is evidence or convention, not yet anyone’s credence. Either way it adds no third basis to the two that are the clinician’s own, and it is those I examine here.
Calling the sedimented sense a credence is not free, and not every tacit state that colors a clinician’s reading earns the name. A practical inclination to act, a perception of danger, a preference among outcomes, a bare feeling of fit—each shapes interpretation without being a degree of belief about what will happen. What makes a tacit state a prior is structure: that it orders the relevant outcomes by expected plausibility, bears on the same question the coming evidence will, and can be drawn out, at least as a range, as something distinct from what the clinician wants or would do. A category judgment or a sense of mechanism may feed such a credence—the type fixes a reference class, the mechanism a direction—but the credence is the ordering of outcomes they yield, not the perception itself.
This marks a boundary the reconstruction must respect, and respecting it is what keeps the tacit basis from being either overclaimed or explained away. As a disposition—what the surgeon expects, is surprised by, would wager on—the sedimented credence is real before anyone elicits it; on the subjectivist view a degree of belief is fixed by such dispositions, not by the ability to state a number, which is exactly why it can be a credence while remaining unsayable. What it is not, until it is drawn out, is a computational prior: a distribution one can place in Bayes’ theorem and update. Elicitation—as an interval, a family of curves, a betting disposition, a calibrated forecast—does not create the credence but measures it, and because a real disposition is only roughly coherent, what it recovers is a band rather than a point (Figure 2). Before that step the sediment is prior-relevant and shapes reasoning; only after it can a unique posterior be computed. The distinction is not tacit versus real but credence-as-disposition versus prior-as-elicited-distribution—and it is the first, the disposition, that the account of tacit knowing below concerns.

Figure 2. Two routes into the same machinery. A reference-class prior is built by taking a base rate from a published cohort and judging which class fits this patient; a sedimented prior begins as a tacit expectation that must be drawn out—elicited as a probability or a range—before it can be updated. Elicitation does not make the expectation real, only sayable. The routes converge on the same Bayesian update but are audited differently: a reference-class prior by the aptness and transportability of its cohort, a sedimented prior by its calibration and the validity of the domain in which it formed.
What the two bases share is a common form; what they differ in is what each answers to, and when that answer is available. A reference-class prior answers to the aptness of its chosen population, contestable at the time of use by asking which class was used and whether it was the right comparison. A sedimented prior cannot usually answer by stating its grounds; it answers chiefly through the outcomes and feedback that formed it, and through whether the present case falls within the domain in which that learning occurred. This is why the ordering can reverse: a reference-class prior can be perfectly explicit, a published number anyone can read, yet drawn from the wrong population, while a sedimented sense, opaque about its grounds, may be the best calibrated of all if it was laid down in a setting of real and repeated feedback.
Put in the discipline’s terms: a reference-class prior admits an internalist audit—its grounds are statable, so scrutiny can address the content itself—while a sedimented prior admits chiefly an externalist one, its appraisal ultimately falling to the reliability of the process that formed it, read off calibration. The mixture is not ad hoc; it follows from access. The mode of appraisal available for a prior is fixed by what its basis makes inspectable, and that is the precise sense in which basis determines accountability.
And no real prior is purely one kind. Sizing up a patient, a surgeon draws at once on the rates from half-remembered trials, a sense of which patients resemble this one, and the residue of a thousand cases. That fusion is the point: the two are not bins that whole priors fall into but ingredients a working prior is compounded from, differing in what would justify or undermine each, and each answerable on its own terms. Nor do the two exhaust the ingredients: mechanistic reasoning, causal models, and testimony can each feed a credence, and each answers to standards of its own—causal adequacy, empirical support, the reliability of the source [24,25]. What I claim for the pair examined here is not exclusivity but centrality: they are the contributions most often fused at the bedside, and the two whose modes of accountability differ most instructively.
A prior’s basis is also independent of its strength—how concentrated or diffuse it is, how far the data are allowed to move it (Figure 3). A reference-class prior may be held loosely or with conviction; a hard-won clinical sense may be tentative or carry the weight of a thousand cases. Strength tells you how far a prior will move the posterior; basis tells you what sort of thing is doing the moving, and whether it can be inspected at all.

Figure 3. This isolates a prior’s strength—how far it lets the data move the posterior—from its basis. The priors share a center; what changes top to bottom is how concentrated they are. Against the same thin study, a weakly informative prior barely commits, so the posterior largely follows the study, while a sharp prior holds it near the prior. No panel is the correct one: the weakest prior is not a neutral default but a positive choice to let the data dominate, sound only where the data are heavy enough to trust; where they are thin, as across much of surgery, following them is deference to noise. Dashed, prior; dotted, study; solid, posterior.
The first default is that the published result is the verdict—the estimate read as the answer, significance read as a yes-or-no. Both mistake an input for a conclusion, and the taxonomy says why twice over: the estimate’s path to this patient runs through a judgment of class, and the verdict’s path to a confidence runs through a prior. A point estimate is a number earned in a particular study population—a rate, a risk difference, a hazard ratio—and it is silent about which of this patient’s many classes should govern its transport; significance reports only how surprising the data would be under the null—its yes-or-no an artifact of a threshold, a verdict the fragility-index literature has shown can flip on a single reclassified outcome [3].
The reference-class prior deserves a closer look, because it is the most easily mistaken for something it is not. It is where two careful surgeons most visibly part ways: one recalls a cohort in which many patients did not heal as well as expected, another a series in which nearly all did, each able to name the study behind the number. Their disagreement is not between kinds of prior but between two reference-class priors anchored to different populations. Which population counts is itself a judgment, not a reading off the data: a 72-year-old diabetic smoker with a distal radius fracture belongs at once to an age cohort, a comorbidity cohort of impaired healers, and a cohort defined by bone quality, each with its own base rate, none uniquely the right comparison. This is the reference-class problem [6,7], whose centrality to the standard model of medical prediction—every step from a study population to a probability for this patient passes through it—has been pressed in general form [30]. It is why this basis is the more treacherous of the two: it looks objective, a published number anyone can read from a table, while resting on a consequential and easily hidden judgment about which class to use.
What takes the verdict’s place is a posterior, and how far the trial moves it depends on the prior it meets—and on how much data there are to move it. Where evidence is abundant, defensible priors are pulled together onto it and their differences wash out; where it is thin, the prior carries much of the inference on its own. Surgery is a mostly small-n literature—only rarely the sample sizes that would swamp a reasonable prior—so for many clinical questions the prior is not a footnote to the evidence but a large part of the answer.
The second default is that the surgeon’s expectation is merely a bias. Concede the premise at once: a sedimented prior is a bias. It is a leaning laid down before the present case, a tendency to expect one outcome over another that the data have not yet earned, and there is no use pretending otherwise. But the same is true of the rest. A base rate is a leaning toward the rate of a chosen class; a reference class is a leaning toward one population among the many a patient belongs to. Bias, in the neutral sense, is simply a leaning, and every prior is one. The error in the debiasing reflex is not that it calls the surgeon’s expectation a bias; it is that it stops there, treating the formal input as the neutral baseline and the expectation as the bias to be disciplined away, while the base rate passes untouched. Once every prior is seen as a leaning, the question is no longer whether the expectation is a bias but whether it is a reasonable one, and that question is put to the gut and to the base rate on the same terms.
Of the two, the sedimented prior is the less auditable and the harder to put into words, and a major strand of 20th-century philosophy explains why. Ryle’s regress against the assumption that intelligent action is applied theory—if applying knowledge required consulting a rule, applying that rule would require a further rule, without end—makes knowing how the more basic kind, not knowing that in compressed form [8]. Polanyi’s tacit dimension—we know more than we can tell—is the same point [9]; Merleau-Ponty grounds it in the body, skilled perception operating ahead of any proposition the agent could state [10]; and the tradition in medical epistemology that takes tacit knowledge to be ineliminable rests on exactly this [11,12]. Together the claims say something stronger than that skill is hard to verbalize: explicit knowledge floats on a tacit base, and the descent ends not at a smaller or vaguer proposition but below statability altogether. The surgeon senses, before any list of risk factors, that this patient is one who will struggle, and the reasons, recited afterward, might never quite contain the sense. A sedimented prior is a confidence laid down by a career that never took the form of a claim, which is why it cannot be audited proposition by proposition. But it still functions as prior credence: it shapes how evidence is interpreted and therefore belongs in any adequate Bayesian reconstruction of clinical judgment.
If the prior cannot be stated, it cannot be audited the way a reference-class prior can: there is no content to inspect, no class to name. Its accountability must run through something other than what it says, and two things remain that it cannot hide from. The first is coherence, read off a single inference: the leaning shows itself in how the surgeon updates—what moves the expectation, and by how much. A prior that refuses to move when the evidence is strong and bears on the question, by more than its own width could license, can be convicted without ever being stated, because the posterior does not follow from the prior and the likelihood the surgeon is working with. But coherence is necessary, not sufficient. A confidence held tightly enough will absorb a strong result and barely shift, and that is coherent—an impeccable update from an unearned certainty. Coherence catches the surgeon who updates wrongly; it cannot catch the one who updates rightly from a prior he was never entitled to.
What catches that one is calibration, read off the series rather than the case. A sedimented prior is warranted to the extent that the expectations a surgeon forms and revises this way, across many patients, track what then happens, and suspect to the extent they do not. This is the reliability of the faculty, taken from its outputs because its contents are sealed. Calibration alone, though, is not enough: a surgeon who assigns every patient the base rate is perfectly calibrated and possesses nothing. A sedimented prior earns its keep by discriminating—the cases it rates worse must fare worse—and, where the claim is that the trained eye adds something beyond the chart, by adding predictive value over what the stated variables already give. It is also where the conditions of formation belong. A regular domain with timely, unambiguous feedback is what allows a tacit prior to become calibrated [13,15]; a low-validity corner of practice—referral patterns that skew the case mix, follow-up that selectively returns, the vivid complication that overweights a rare event—is where one should expect noise wearing the costume of judgment [14]. Those conditions are an indicator of whether calibration is possible, not the test; the test is the calibration itself.
With both in hand, the reversal is no longer surprising. A feedback-formed expectation and a base rate are both leanings, and judged by the same two measures the tacit one can win—when it is coherent and calibrated, and the base rate, however large the study behind it, is drawn from a population this patient does not belong to. What should move a prior is not the size of a trial but the weight of evidence that bears on the question; a large study in the wrong population is weak evidence, and declining to move on it is sense, not rigidity. Where that holds, the discipline belongs on the formal input that never faced an outcome, not on the judgment that did. So the section’s claim can be put exactly: the expectation is not merely the bias—not the one the reflex singles out to suppress while the number goes free. It is a leaning, as every prior is, including the tacit ones that cannot be quietly cleared out of a subjectivism otherwise accepted. Whether it is the distorting kind is settled by whether it coheres, calibrates, discriminates, and is being used within the domain where it formed—not by its being a leaning, and not by its being tacit.
To represent a sedimented prior in Bayesian terms is not to claim the surgeon carries a defined distribution in his or her head and multiplies it by the data giving a mathematically defined posterior. The Bayesian account here is a rational reconstruction of the reasoning structure, not a literal description of cognition, a way of making coherent the commitments a tacit judgment leaves silent. Its warrant is the coherence and calibration the reconstruction makes checkable, not the resemblance between a prior and the phenomenology of skill. Recent evidence suggests that clinicians reason in mixed rather than purely Bayesian ways [16], consistent with the broader literature on departures from Bayesian norms [14].
The third default is that disagreement between careful readers is error. What replaces it is not license: none of this implies that any prior is as good as another. The common misreading of subjectivists like de Finetti is that subjective means arbitrary. A subjective prior is still constrained: it must be internally consistent and it must respect the clinician’s own commitments [17]. Coherence rules out obvious contradictions such as violating an axiom of probability. And once the clinician accepts trials as evidence, coherence requires updating on them. Respecting the literature is therefore not an external rule; it follows from what the clinician already believes. Still, critics are right that an unconstrained subjectivism would permit beliefs we would normally call unreasonable [6,18]. Priors must remain answerable to mechanism, to the literature, and to physiological plausibility. The notion can be made precise. A prior lies within the defensible region when it is coherent with the agent’s other credences, responsive to the total evidence the agent accepts as relevant, and supportable under publicly contestable epistemic norms: calibration, causal adequacy, and appropriate reference-class selection among them. So defined, the region is neither wholly objective nor merely personal: its boundaries depend in part on an agent’s evidential situation, yet remain open to debate, because the norms that fix them are public rather than private. This is the constrained subjectivism I defend: subjectivist because the prior is a degree of belief answerable to the agent’s own commitments, constrained because those commitments, and the norms governing them, are not arbitrary. More than one prior survives these constraints and more than one fails: the region rules many beliefs out without narrowing to a single correct one, so the view is neither objective Bayesianism nor permissive relativism. That competent surgeons sometimes disagree within these bounds rather than outside them is what the view predicts, not a defect in it. Such disagreement has two sources the taxonomy locates, and they are not the same phenomenon. Two surgeons may anchor to different but defensible reference classes—the two cohorts of the earlier case—and this is genuine slack within shared evidence: one body of data, more than one rational reading, because the data do not close the choice of class. Or they may carry different sedimented priors laid down by different careers—and this is not slack in shared evidence at all, but divergence from different total evidence, each career being an evidential history the other surgeon does not have. Neither is a failure to reason, but only the first should trouble anyone who thinks evidence dictates a unique response.
The distinction earns its keep against a near neighbor in general epistemology. Whether a single body of evidence can rationally support more than one doxastic attitude is the debate over permissivism [31–33], and the reference-class slack is, in that vocabulary, a permissive case—while the sedimented divergence is not, since divergence under different evidence is one every party to that debate already allows. I do not take a side in the general debate here; the claim is narrower. Where clinical disagreement is permissive at all, its permission has a locatable source—a choice of class the evidence cannot close—and contestable bounds, the public norms just named, which is what keeps the concession from collapsing into the arbitrariness that critics of permissivism fear [31].
The boundary of the defensible region is neither fixed nor timeless. Two orthopaedic surgical examples illustrate this. Arthroscopic debridement and partial meniscectomy for the osteoarthritic knee once rested on an impeccable mechanistic prior: a degenerative tear is a structural lesion, and resecting it ought to relieve the pain it causes. Yet meniscal tears turn up in most asymptomatic age-matched knees, their presence barely tracking pain or function [19], and the sham-controlled trials that followed dismantled the rationale [20,21]. Vertebroplasty for osteoporotic vertebral fracture traveled the same arc, from a near-self-evident mechanical rationale to refutation under blinded sham comparison [22,23].
The decisive point is not that these priors were irrational: in their moment they were defensible, held by competent clinicians reasoning from the best mechanism and evidence then available [24,25], but that the defensible region moved (Figure 4). No prior chosen once and fixed in advance stays defensible on its own; what keeps it honest is not where it came from but its continued answerability: to mechanism, to the literature, and above all to the evidence still to come. The original confidence rested on mechanism and on the high improvement rates reported by early studies without a control group. What overturned it was not fresh data defeating a settled belief so much as a change in the comparison that mattered: not patients before and after treatment, but patients given the procedure rather than a convincing sham. The better comparison changed which observed rate was the apt one to carry forward to the next patient, and which rate is apt is itself a reference-class judgment. Even the data leave open which comparison counts, and that choice is the same kind of judgment, one basis down. The objective-Bayesian program, which constructs formal or default priors to solve real statistical problems [4,5,26,27], is untouched by these reversals; such a prior is a methodological construct, not a credence a clinician holds. What the reversals show is narrower, and enough here: the warrant of any prior rests on its continued answerability rather than on where it came from.

Figure 4. The defensible region is bounded but not fixed. A prior can sit within it and later fall outside as evidence accumulates, without thereby having been irrational when first held; the region itself answers to the data. The marker tracks a pro-procedure prior before and after sham-controlled trials.
Objection 1: the taxonomy classifies causal origins, not epistemic warrants. One might grant that class and sediment name where a credence comes from while denying that origin bears on justification, since genetic and epistemic questions are rightly kept apart. But origin here is not doing genetic work. A reference class can be checked for aptness against recorded outcomes; a sediment, if at all, only through the reliability of its outputs. The basis does not confer the warrant; it fixes what would count as scrutinizing the credence at all.
Objection 2: mixed bases make source-sensitive scrutiny impossible. If every actual prior is a composite, the source-specific standards seem to have nothing determinate to apply to. But the argument requires load-bearing components, not clean decomposition. Faced with a confident prior, one asks what is carrying the weight: a class whose rate supports it, or an impression reaching for a citation after the fact? Whichever basis dominates can be identified and questioned on its own terms, even with the others present.
Objection 3: dispositional reconstruction does not determine a unique credence. This is the deepest objection, and it is correct. A tacit sense shows itself in what a surgeon expects, what surprises them, and what they choose, and that behavior does not single out one exact number. But the conclusion is not that the prior is a fiction; it is that expecting a single number was the mistake. The honest representation is a range, as wide as the behavior leaves open and no wider, which updates on evidence as an ordinary prior does but with the whole band shifting at once [18,28,34]—the confidence with no sharp edges that the blurred panel of Figure 2 already depicts.
A prior earns its place in clinical reasoning by being answerable, but what it is answerable to—and when—depends on what kind of thing it is. Three habits of reading follow, each overturning a default: the published result is an input, not a verdict; the surgeon’s expectation can be a calibrated prior, not merely a bias; and disagreement between careful readers can mark two defensible priors rather than an error to be stamped out. A clinical prior, in the end, may rest on more than one epistemic basis: a base rate or the sediment of experience, neither peculiar to medicine but the record and the residue that every practical judgment draws on. The Bayesian apparatus is less a way of computing with the prior than a way of being direct and transparent about it. Read this way, the theorem is not a calculator the surgeon runs but a discipline of disclosure: say what you believed before the trial, say where that belief came from, and stay answerable for it once the evidence arrives. The humility it asks is a principled refusal to mistake one defensible reading of a trial for the only one, which is, after all, the difference between holding a prior and not knowing you have one.
[1] de Finetti B. Theory of Probability: A Critical Introductory Treatment. Vol 1. New York: Wiley; 1974.
[2] Goodman SN. Toward evidence-based medical statistics. 1: The P value fallacy. Ann Intern Med. 1999;130(12):995–1004.
[3] Walsh M, Srinathan SK, McAuley DF, et al. The statistical significance of randomized controlled trial results is frequently fragile: a case for a Fragility Index. J Clin Epidemiol. 2014;67(6):622–628.
[4] Jeffreys H. Theory of Probability. 3rd ed. Oxford: Clarendon Press; 1961.
[5] Jaynes ET. Probability Theory: The Logic of Science. Cambridge: Cambridge University Press; 2003.
[6] Hájek A. The reference class problem is your problem too. Synthese. 2007;156(3):563–585.
[7] Cartwright N, Hardie J. Evidence-Based Policy: A Practical Guide to Doing It Better. Oxford University Press; 2012.
[8] Ryle G. The Concept of Mind. London: Hutchinson; 1949.
[9] Polanyi M. The Tacit Dimension. London: Routledge & Kegan Paul; 1966.
[10] Merleau-Ponty M. Phenomenology of Perception. Smith C, trans. London: Routledge & Kegan Paul; 1962. (Original work published 1945.)
[11] Thornton T. Tacit knowledge as the unifying factor in evidence based medicine and clinical judgement. Philos Ethics Humanit Med. 2006 Mar 17;1(1):E2. doi: 10.1186/1747-5341-1-2. PMID: 16759426; PMCID: PMC1475611.
[12] Henry SG. Recognizing tacit knowledge in medical epistemology. Theoretical Medicine and Bioethics. 2006;27(3):187–213.
[13] Kahneman D, Klein G. Conditions for intuitive expertise: a failure to disagree. Am Psychol. 2009;64(6):515–526.
[14] Kahneman D, Tversky A. On the psychology of prediction. Psychol Rev. 1973;80(4):237–251.
[15] Gigerenzer G, Brighton H. Homo heuristicus: why biased minds make better inferences. Top Cogn Sci. 2009;1(1):107–143.
[16] Parisien R, Drost A, Razi A, Ramtin S, Ring D, Janssen SJ (2026) Musculoskeletal surgeons use mixed reasoning rather than pure Bayesian strategies in clinical practice. PLoS One 21(6): e0351694.
[17] Lin H. Bayesian epistemology. In: Zalta EN, Nodelman U, eds. The Stanford Encyclopedia of Philosophy. 2022.
[18] Eriksson L, Hájek A. What are degrees of belief? Studia Logica. 2007;86(2):183–213.
[19] Bhattacharyya T, Gale D, Dewire P, et al. The clinical importance of meniscal tears demonstrated by magnetic resonance imaging in osteoarthritis of the knee. J Bone Joint Surg Am. 2003;85(1):4–9.
[20] Moseley JB, O’Malley K, Petersen NJ, et al. A controlled trial of arthroscopic surgery for osteoarthritis of the knee. N Engl J Med. 2002;347(2):81–88.
[21] Sihvonen R, Paavola M, Malmivaara A, et al. Arthroscopic partial meniscectomy versus sham surgery for a degenerative meniscal tear. N Engl J Med. 2013;369(26):2515–2524.
[22] Buchbinder R, Osborne RH, Ebeling PR, et al. A randomized trial of vertebroplasty for painful osteoporotic vertebral fractures. N Engl J Med. 2009;361(6):557–568.
[23] Kallmes DF, Comstock BA, Heagerty PJ, et al. A randomized trial of vertebroplasty for osteoporotic spinal fractures. N Engl J Med. 2009;361(6):569–579.
[24] Howick J. The Philosophy of Evidence-Based Medicine. Wiley-Blackwell; 2011.
[25] Russo F, Williamson J. Interpreting causality in the health sciences. International Studies in the Philosophy of Science. 2007;21(2):157–170.
[26] Bernardo JM. Reference posterior distributions for Bayesian inference. J R Stat Soc Series B. 1979;41(2):113–147.
[27] Berger J. The case for objective Bayesian analysis. Bayesian Anal. 2006;1(3):385–402.
[28] Walley P. Statistical Reasoning with Imprecise Probabilities. London: Chapman & Hall; 1991.
[29] Stegenga J. Medical Nihilism. Oxford: Oxford University Press; 2018.
[30] Fuller J, Flores LJ. The Risk GP Model: the standard model of prediction in medicine. Stud Hist Philos Biol Biomed Sci. 2015;54:49–61.
[31] White R. Epistemic permissiveness. Philosophical Perspectives. 2005;19(1):445–459.
[32] Schoenfield M. Permission to believe: why permissivism is true and what it tells us about irrelevant influences on belief. Noûs. 2014;48(2):193–218.
[33] Kelly T. Evidence can be permissive. In: Steup M, Turri J, Sosa E, eds. Contemporary Debates in Epistemology. 2nd ed. Chichester: Wiley-Blackwell; 2013:298–312.
[34] Joyce JM. A defense of imprecise credences in inference and decision making. Philosophical Perspectives. 2010;24(1):281–323.