My research asks how we know what we claim to know in clinical practice: how surgeons reason under uncertainty, and where our confidence is and is not warranted. Much of that work is linguistic, since the assumptions are usually carried in how a claim is worded rather than stated outright. These are questions the specialty has had comparatively little occasion to take up, and I am interested in what happens when we do.
Philosophical inquiry can (I think) help us examine what our measurements are measurements of, how categories acquire boundaries, and when evidence truly warrants a conclusion. That work seems more pressing now that our tools for interpreting the literature are changing faster than the concepts underneath it. This is a small attempt to clarify the language we rely on, and to keep what we know, and how we justify that, in view while the tools change around us.
Orthopaedic surgeons have not converged on a precise definition of fracture union in over twenty years. This project suggests why that is hard: healing is continuous, and where the line falls between united and not united is a convention adopted for a purpose, not a boundary waiting to be found. In a corpus of 296 research articles, 214 defined the outcome, and 200 of those defined it by a rule someone applies. Of those 200, 92% went on to report the result in the language of discovery: union occurred, nonunion developed. The criterion is usually stated plainly in the Methods; it just does not travel with the number. And the number is what gets pooled, compared, and inherited by the next study.
Every clinical encounter is dense with unacknowledged philosophical commitments about knowledge, causation, probability, and value. Through a routine clavicle fracture encounter, this article identifies five philosophical concepts embedded in daily orthopaedic practice, from theory-laden observation to the reference class problem. The through-line is epistemic humility: a disciplined awareness of the boundaries of what can be known, and a willingness to hold our assumptions a bit less firmly, a safeguard against the complacency of the unexamined practice.
Abstracts and discussion sections of the same paper are not epistemically equivalent, yet readers and AI systems treat them as if they are. In a corpus analysis of 201 publications from JBJS and CORR, this project quantifies the divergence between the confident language of abstracts and the hedged, qualified language of full-text discussions. The gap is systematic, measurable, and has direct consequences for how evidence is interpreted, synthesized, and consumed by large language models trained on abstract-heavy data.
When deciding whether to adopt a new surgical innovation, the evidence underdetermines the choice because of the reference class problem, usually posed about patients but here applied to the surgeon. Trial evidence rarely tells us whether the conditions under which it was generated, including surgeon experience, volume and institutional support, match our own. A surgeon's experience, learning curve, and assessment of potential harms are unique, so a decision to adopt can be helpful under one set of circumstances and harmful under another. The issue is fundamentally an epistemological one: how do we know if we should use a new technique?
Explore the interactive model at RCPmodel.netlify.app →A Bayesian re-analysis of the CROSSFIRE distal radius fracture trial, showing that defensible surgeon readings of the same data can yield probabilities of meaningful benefit ranging from essentially zero to 0.47 at three months. The divergence is a structural feature of how clinical trials are read, emphasizing the importance of prior conditionals. The underlying claim is that there are no prior-free readings of any evidence: priors are subjective, which allows two surgeons to disagree and yet both be reasoning well.
Bayesian probability is notoriously difficult to apply consistently in clinical settings. It remains unclear how often surgeons actually utilize that reasoning strategy despite tacitly recognizing its normative value. In this survey of 153 surgeons reasoning through eight scenarios of test and treatment decisions, most showed mixed patterns, acknowledging prior probability but underweighting it, without explicit updating, and reasoning strategy varied with clinical context rather than forming a unified style. Bayesian updating is the normative ideal, but actual clinical reasoning is context-dependent and heuristic-laden; the gap between the two may be a matter of trainable skill rather than fixed disposition.
Cognitive biases are ubiquitous in human reasoning, and orthopaedic decision-making is no exception. This study evaluated the prevalence of specific biases, including anchoring, availability, and framing effects, among academic orthopaedic surgeons. What is at stake is how we come to know what we think we know, what counts as a justified belief, and whether our beliefs are genuinely updatable in the face of new evidence.
Much of orthopaedic decision-making rests on uncertain evidence, yet this survey of 242 surgeons found strikingly low recognition of uncertainty, and it did not improve with years in practice, while overconfidence grew. Less recognition of uncertainty was associated with greater confidence bias and greater trust in the evidence base; better statistical understanding was the strongest predictor of acknowledging uncertainty. Our confidence may be less well calibrated to what the evidence can support than we assume, and the findings suggest calibration is a learnable skill rather than an automatic byproduct of experience.
As artificial intelligence increasingly surpasses clinicians in predictive accuracy, defending the human role on superior intuition or diagnostic skill becomes harder to sustain. This paper argues that every clinical prediction is a conditional: given an antecedent, the chosen endpoint, reference class, and weighting of harms, the system supplies a consequent, the predicted outcome or recommended action. Accuracy on the consequent does not settle whether the antecedent is right for this patient. The clinician’s enduring role is therefore not to supervise the algorithm’s prediction, but to adopt that frame for the individual patient and remain answerable for the act that follows. That is the basis for the paper’s “centaur”: a human–machine partnership in which the machine predicts and the clinician answers for how that prediction is made clinically operative.
Submitted September 2026 to Philosophy & Technology, collection “The ethics of medical artificial intelligence: trust and adoption of AI in healthcare”
The Duhem-Quine problem is usually presented from the failure side: a result that contradicts the hypothesis indicts the conjunction without settling how blame is distributed. Success carries a similar underdetermination, since a confirming result supports the conjunction without settling how credit is distributed. Clinical research responds vigorously to the first problem and far less to the second, and I call this allocation of scrutiny the audit asymmetry. It is institutional rather than individual, and its costs are greatest where published success dominates, practice is costly to reverse, and trials carry heavy auxiliary loads, as in surgery. Where a field's priors rest on its successes, unaudited credit hardens into the expectations that shield established practice, and the Duhem-Quine problem, invoked in its defense, applies equally to the successes that established it. The remedy asks more than epistemic humility: it asks a field to put to its successes the question it already puts to failure.
Submitted August 2026 to The Journal of Medicine & Philosophy
Fracture studies often define nonunion by what a surgeon did about it: a return to the operating room to promote healing. The definition is stated in the Methods. The result is then reported as a fact about the bone, and the same operation that defined the outcome is presented as the thing the outcome made necessary. When the decision to operate is the evidence that the operation was needed, the argument is circular, and the rate it produces is a rate of decisions rather than a rate of biology. This study counts how often that happens in recent papers from four leading orthopaedic journals: how often a treatment decision contributes to the nonunion endpoint, how often that dependence survives into the abstract, and how often the paper goes on to draw a conclusion about healing from it. The questions, coding rules, and predictions are registered before the data are read. The study is the clinical counterpart of “Judgment as Discovery” (Journal of Evaluation in Clinical Practice, 2026), which examined the same endpoint from the language side.
In progress: targeting Clinical Orthopaedics and Related Research®, 2027
When two competent physicians examine the same trial data for the same patient, can they radically disagree and both still be rational? Modern medicine relies heavily on probabilistic risk estimates, a framework that mathematically requires prior probabilities. Yet applying general population data to a single, complex patient inevitably forces judgment calls about which evidence matters most, making a single “uniquely rational” answer structurally impossible. Taking medicine’s own Bayesian logic seriously, this paper argues that practice variation is not always a scandal of bad reasoning or bias. Instead it argues for constrained permissivism: a framework in which evidence naturally permits a bounded range of defensible conclusions, offering a rigorous epistemological foundation for why rational experts legitimately disagree.
In progress: targeting Theoretical Medicine and Bioethics, Philosophy of Medicine, or the Journal of Medicine and Philosophy.
A scoping review of the RUST, mRUST and RUSH score families, currently in progress. The scores did what they were built to do, since readers agree with each other more than with unaided impression, but the threshold on the score has continued to move: between reports of the same fracture at the same site, with the implant, and within a single cohort with the follow-up interval at which it is measured. The review asks three questions of the literature rather than of anyone’s practice: where thresholds have been set and whether their dispersion has narrowed; whether an adopted threshold still travels with the site, fixation, interval and reference standard it was derived from; and whether the standards those thresholds were calibrated against coincide at all. Two kinds of threshold turn out to be in use, one certifying that union is present and one predicting failure not yet arrived, carrying opposite costs of error, so a difference between them is structure rather than disagreement.
In progress…
« Le mieux est l’ennemi du bien. » — Voltaire
When is one more maneuver worth it? An expected-utility exploration of the stable but imperfect proximal humerus construct. One more attempt has positive expected value only when
None of these quantities can be read off the fluoroscopy image; each is set by surgeon-specific priors. Defensible priors generate a band of rational stopping points, not a single answer, and the final radiograph records where a surgeon stopped, not whether stopping there was wise.
Bayesian talk is everywhere in clinical reasoning, but the prior itself is rarely examined. Is it a summary of past frequencies, a personal degree of belief, a commitment answerable to argument, or some shifting mix? Why getting the answer wrong leads surgeons to treat contestable starting points as settled facts.
A census of how the orthopaedic literature handles the assumptions it starts from. Priors are necessarily subjective, yet the subjective Bayesian view, in the tradition of de Finetti, is not the one taken by the overwhelming majority of papers currently being written. Meanwhile "Bayes" and "Bayesian" as search terms have been increasing steadily in frequency, especially over the past five years: the vocabulary is spreading faster than the philosophical commitments it carries.
A planned study of Introduction sections, currently at concept stage. The Introduction is the only part of a paper that makes empirical claims without presenting evidence: a finding that was a hedged association in a retrospective cohort reappears, two citations later, as a bare declarative with a superscript. Nothing in that sentence is false, and every existing study of citation accuracy would score it as correct; the loss is in the grammar rather than the content. The protocol codes what kind of evidence a citing sentence discloses, and, where the source can be retrieved, whether certainty, scope, or causal strength has increased in the retelling. Compression is demanded by the form, so the aim is a reporting convention rather than a charge of carelessness.
High-frequency, single-item functional check-ins by text message build continuous recovery curves for individual patients: improvement velocity, total disability burden, and the point at which recovery plateaus. A move from asking whether the bone united to whether the patient recovered.