My research asks how we know what we claim to know in clinical practice: how surgeons reason under uncertainty, and where our confidence is and is not warranted. Much of that work is linguistic, since the assumptions are usually carried in how a claim is worded rather than stated outright. These are questions the specialty has had comparatively little occasion to take up, and I am interested in what happens when we do.
Philosophical inquiry can (I think) help us examine what our measurements are measurements of, how categories acquire boundaries, and when evidence truly warrants a conclusion. That work seems more pressing now that our tools for interpreting the literature are changing faster than the concepts underneath it. This is a small attempt to clarify the language we rely on, and to keep what we know, and how we justify that, in view while the tools change around us.
Every clinical encounter is dense with unacknowledged philosophical commitments about knowledge, causation, probability, and value. Through a routine clavicle fracture encounter, this article identifies five philosophical concepts embedded in daily orthopaedic practice, from theory-laden observation to the reference class problem. The through-line is epistemic humility: a disciplined awareness of the boundaries of what can be known, and a willingness to hold our assumptions a bit less firmly, a safeguard against the complacency of the unexamined practice.
Abstracts and discussion sections of the same paper are not epistemically equivalent, yet readers and AI systems treat them as if they are. In a corpus analysis of 201 publications from JBJS and CORR, this project quantifies the divergence between the confident language of abstracts and the hedged, qualified language of full-text discussions. The gap is systematic, measurable, and has direct consequences for how evidence is interpreted, synthesized, and consumed by large language models trained on abstract-heavy data.
Orthopaedic surgeons have not converged on a precise definition of fracture union in over twenty years. This project suggests why that is hard: healing is continuous, and where the line falls between united and not united is a convention adopted for a purpose, not a boundary waiting to be found. In a corpus of 296 research articles, 214 defined the outcome, and 200 of those defined it by a rule someone applies. Of those 200, 92% went on to report the result in the language of discovery: union occurred, nonunion developed. The criterion is usually stated plainly in the Methods; it just does not travel with the number. And the number is what gets pooled, compared, and inherited by the next study.
Much of orthopaedic decision-making rests on uncertain evidence, yet this survey of 242 surgeons found strikingly low recognition of uncertainty, and it did not improve with years in practice, while overconfidence grew. Less recognition of uncertainty was associated with greater confidence bias and greater trust in the evidence base; better statistical understanding was the strongest predictor of acknowledging uncertainty. Our confidence may be less well calibrated to what the evidence can support than we assume, and the findings suggest calibration is a learnable skill rather than an automatic byproduct of experience.
Cognitive biases are ubiquitous in human reasoning, and orthopaedic decision-making is no exception. This study evaluated the prevalence of specific biases, including anchoring, availability, and framing effects, among academic orthopaedic surgeons. What is at stake is how we come to know what we think we know, what counts as a justified belief, and whether our beliefs are genuinely updatable in the face of new evidence.
Bayesian probability is notoriously difficult to apply consistently in clinical settings. It remains unclear how often surgeons actually utilize that reasoning strategy despite tacitly recognizing its normative value. In this survey of 153 surgeons reasoning through eight scenarios of test and treatment decisions, most showed mixed patterns, acknowledging prior probability but underweighting it, without explicit updating, and reasoning strategy varied with clinical context rather than forming a unified style. Bayesian updating is the normative ideal, but actual clinical reasoning is context-dependent and heuristic-laden; the gap between the two may be a matter of trainable skill rather than fixed disposition.
When deciding whether to adopt a new surgical innovation, the evidence underdetermines the choice because of the reference class problem, usually posed about patients but here applied to the surgeon. Trial evidence rarely tells us whether the conditions under which it was generated, including surgeon experience, volume and institutional support, match our own. A surgeon's experience, learning curve, and assessment of potential harms are unique, so a decision to adopt can be helpful under one set of circumstances and harmful under another. The issue is fundamentally an epistemological one: how do we know if we should use a new technique?
Explore the interactive model →A Bayesian re-analysis of the CROSSFIRE distal radius fracture trial, showing that defensible surgeon readings of the same data can yield probabilities of meaningful benefit ranging from essentially zero to 0.47 at three months. The divergence is a structural feature of how clinical trials are read, emphasizing the importance of prior conditionals. The underlying claim is that there are no prior-free readings of any evidence: priors are subjective, which allows two surgeons to disagree and yet both be reasoning well.
In clinical orthopaedic trauma research, nonunion is routinely reported as an objective biological failure, yet in practice the research endpoint is frequently constituted by a surgeon’s decision to return to the operating room. This project is a systematic, empirical meta-research study across major orthopaedic trauma journals (JOT, JBJS, CORR) evaluating how often therapeutic reoperation is treated as necessary or sufficient to classify an index fracture as a nonunion. By analyzing how clinical decisions under uncertainty become hard research endpoints, and how this dependence is often masked between the Methods and the Abstract, the study aims to quantify treatment-entangled reporting and establish methodological standards that keep clinical judgment transparent when patient care becomes scientific data.
In progress…
A scoping review of the RUST, mRUST and RUSH score families, currently in progress. The scores did what they were built to do, since readers agree with each other more than with unaided impression, but the threshold on the score has continued to move: between reports of the same fracture at the same site, with the implant, and within a single cohort with the follow-up interval at which it is measured. The review asks three questions of the literature rather than of anyone’s practice: where thresholds have been set and whether their dispersion has narrowed; whether an adopted threshold still travels with the site, fixation, interval and reference standard it was derived from; and whether the standards those thresholds were calibrated against coincide at all. Two kinds of threshold turn out to be in use, one certifying that union is present and one predicting failure not yet arrived, carrying opposite costs of error, so a difference between them is structure rather than disagreement.
In progress…
The Duhem-Quine problem is usually presented from the failure side: a result that contradicts the hypothesis indicts the conjunction without settling how blame is distributed. Success carries a similar underdetermination, since a confirming result supports the conjunction without settling how credit is distributed. Clinical research responds vigorously to the first problem and far less to the second, and I call this allocation of scrutiny the audit asymmetry. It is institutional rather than individual, and its costs are greatest where published success dominates, practice is costly to reverse, and trials carry heavy auxiliary loads, as in surgery. Where a field's priors rest on its successes, unaudited credit hardens into the expectations that shield established practice, and the Duhem-Quine problem, invoked in its defense, applies equally to the successes that established it. The remedy asks more than epistemic humility: it asks a field to put to its successes the question it already puts to failure.
Submitted 30 August 2026 to The Journal of Medicine & Philosophy
When two competent physicians examine the same trial data for the same patient, can they radically disagree and both still be rational? Modern medicine relies heavily on probabilistic risk estimates, a framework that mathematically requires prior probabilities. Yet applying general population data to a single, complex patient inevitably forces judgment calls about which evidence matters most, making a single “uniquely rational” answer structurally impossible. Taking medicine’s own Bayesian logic seriously, this paper argues that practice variation is not always a scandal of bad reasoning or bias. Instead it argues for constrained permissivism: a framework in which evidence naturally permits a bounded range of defensible conclusions, offering a rigorous epistemological foundation for why rational experts legitimately disagree.
In progress: targeting Theoretical Medicine and Bioethics, Philosophy of Medicine, or the Journal of Medicine and Philosophy.
« Le mieux est l’ennemi du bien. » — Voltaire
When is one more maneuver worth it? An expected-utility exploration of the stable but imperfect proximal humerus construct. One more attempt has positive expected value only when
None of these quantities can be read off the fluoroscopy image; each is set by surgeon-specific priors. Defensible priors generate a band of rational stopping points, not a single answer, and the final radiograph records where a surgeon stopped, not whether stopping there was wise.
Bayesian talk is everywhere in clinical reasoning, but the prior itself is rarely examined. Is it a summary of past frequencies, a personal degree of belief, a commitment answerable to argument, or some shifting mix? Why getting the answer wrong leads surgeons to treat contestable starting points as settled facts.
A census of how the orthopaedic literature handles the assumptions it starts from. Priors are necessarily subjective, yet the subjective Bayesian view, in the tradition of de Finetti, is not the one taken by the overwhelming majority of papers currently being written. Meanwhile "Bayes" and "Bayesian" as search terms have been increasing steadily in frequency, especially over the past five years: the vocabulary is spreading faster than the philosophical commitments it carries.
LLMs are usually benchmarked against expert humans, right or wrong relative to what a good clinician would say. This asks whether that’s the wrong yardstick, testing whether AI tracks the strength and grade of clinical recommendations rather than merely their conclusions.
A planned study of Introduction sections, currently at concept stage. The Introduction is the only part of a paper that makes empirical claims without presenting evidence: a finding that was a hedged association in a retrospective cohort reappears, two citations later, as a bare declarative with a superscript. Nothing in that sentence is false, and every existing study of citation accuracy would score it as correct; the loss is in the grammar rather than the content. The protocol codes what kind of evidence a citing sentence discloses, and, where the source can be retrieved, whether certainty, scope, or causal strength has increased in the retelling. Compression is demanded by the form, so the aim is a reporting convention rather than a charge of carelessness.
High-frequency, single-item functional check-ins by text message build continuous recovery curves for individual patients: improvement velocity, total disability burden, and the point at which recovery plateaus. A move from asking whether the bone united to whether the patient recovered.
An annotated catalogue of exemplar papers in ten themes, each with a short analysis and my own opinion set apart from it.