My research asks how we know what we claim to know in clinical practice: how surgeons reason under uncertainty, and where our confidence is and is not warranted. Much of that work is linguistic, since the assumptions are usually carried in how a claim is worded rather than stated outright. These are questions the specialty has had comparatively little occasion to take up, and I am interested in what happens when they are put to it seriously.
This work has grown out of a preoccupation with what we can actually know in clinical practice, and with the assumptions we make without noticing. What we treat as observation often turns out to be interpretation, and what we treat as measurement often turns out to be a decision about where to draw a line. Surgery is a useful place to examine this, because the decisions are consequential and the evidence is rarely as decisive as the confidence with which we act on it.
Every clinical encounter is dense with unacknowledged philosophical commitments — about knowledge, causation, probability, and value. Through a routine clavicle fracture encounter, this article identifies five philosophical concepts embedded in daily orthopaedic practice, from theory-laden observation to the reference class problem. The through-line is epistemic humility: a disciplined awareness of the boundaries of what can be known, and a willingness to hold our assumptions a bit less firmly — a safeguard against the complacency of the unexamined practice.
Much of orthopaedic decision-making rests on uncertain evidence, yet this survey of 242 surgeons found strikingly low recognition of uncertainty — and it did not improve with years in practice, while overconfidence grew. Less recognition of uncertainty was associated with greater confidence bias and greater trust in the evidence base; better statistical understanding was the strongest predictor of acknowledging uncertainty. Our confidence may be less well calibrated to what the evidence can support than we assume, and the findings suggest calibration is a learnable skill rather than an automatic byproduct of experience.
Cognitive biases are ubiquitous in human reasoning, and orthopaedic decision-making is no exception. This study evaluated the prevalence of specific biases — including anchoring, availability, and framing effects — among academic orthopaedic surgeons. What is at stake is how we come to know what we think we know, what counts as a justified belief, and whether our beliefs are genuinely updatable in the face of new evidence.
Bayesian probability is notoriously difficult to apply consistently in clinical settings. It remains unclear how often surgeons actually utilize Bayesian reasoning despite tacitly recognizing its normative value. In this survey of 153 surgeons reasoning through eight scenarios of test and treatment decisions, most showed mixed patterns — acknowledging prior probability but underweighting it, without explicit updating — and reasoning strategy varied with clinical context rather than forming a unified style. Bayesian updating is the normative ideal, but actual clinical reasoning is context-dependent and heuristic-laden — the gap between the two may be a matter of trainable skill rather than fixed disposition.
Abstracts and discussion sections of the same paper are not epistemically equivalent — yet readers and AI systems treat them as if they are. In a corpus analysis of 201 publications from JBJS and CORR, this project quantifies the divergence between the confident language of abstracts and the hedged, qualified language of full-text discussions. The gap is systematic, measurable, and has direct consequences for how evidence is interpreted, synthesized, and consumed by large language models trained on abstract-heavy data.
When deciding whether to adopt a new surgical innovation, the evidence underdetermines the choice because of the reference class problem — usually posed about patients, but here applied to the surgeon. Trial evidence rarely tells us whether the conditions under which it was generated — surgeon experience, volume, institutional support — match our own. A surgeon's experience, learning curve, and assessment of potential harms are unique — so a decision to adopt can be helpful under one set of circumstances and harmful under another. The issue is fundamentally an epistemological one: how do we know if we should use a new technique?
Explore the interactive model →A Bayesian re-analysis of the CROSSFIRE distal radius fracture trial, showing that defensible surgeon readings of the same data can yield probabilities of meaningful benefit ranging from essentially zero to 0.47 at three months. The divergence is a structural feature of how clinical trials are read — emphasizing the importance of prior conditionals. The underlying claim is that there are no prior-free readings of any evidence: priors are subjective, which allows two surgeons to disagree and yet both be reasoning well.
For over two decades, orthopaedic surgeons have not coalesced on a precise definition of fracture union. This project suggests the impasse may be a category error rather than a coordination problem: union is a clinical judgment imposed on a continuous biological process, not a biological state awaiting discovery. A corpus analysis of 300 research articles finds that 93% define union as a decision — then report it in the language of discovery. That discovery language is a linguistic phenomenon rather than necessarily a sign of intent, but the move has consequences: it complicates comparison between studies and shapes how the field regards the very possibility of consensus.
A tibia at fourteen weeks: three cortices bridged, the fourth still showing a line, and two experienced surgeons who disagree. Our usual vocabulary for that — observer variability, measurement error, inadequate standardisation — assumes somebody is mistaken. This raises a fourth possibility: that nothing has gone wrong, and the disagreement belongs to the question. United is a vague predicate and behaves as heap does, with a penumbra of borderline cases where careful people sort identical images differently. Three consequences follow. Published thresholds sitting at 10, 11, 12 and 13 need not mean anyone is wrong; sharpening the scale moves the boundary rather than removing it; and unlike noise, vagueness predicts that disagreement will concentrate at the boundary rather than spread across the range. The practical suggestion is modest: say what your threshold is, say what it is for, and report the rate at neighbouring values.
Draft available on request.
A scoping review of the RUST, mRUST and RUSH score families, currently in progress. The scores did what they were built to do — readers agree with each other more than with unaided impression — but the threshold on the score has continued to move: between reports of the same fracture at the same site, with the implant, and within a single cohort with the follow-up interval at which it is measured. The review asks three questions of the literature rather than of anyone’s practice: where thresholds have been set and whether their dispersion has narrowed; whether an adopted threshold still travels with the site, fixation, interval and reference standard it was derived from; and whether the standards those thresholds were calibrated against coincide at all. Two kinds of threshold turn out to be in use — one certifying that union is present, one predicting failure not yet arrived — carrying opposite costs of error, so a difference between them is structure rather than disagreement.
A planned study of Introduction sections, currently at concept stage. The Introduction is the only part of a paper that makes empirical claims without presenting evidence: a finding that was a hedged association in a retrospective cohort reappears, two citations later, as a bare declarative with a superscript. Nothing in that sentence is false, and every existing study of citation accuracy would score it as correct — the loss is in the grammar rather than the content. The protocol codes what kind of evidence a citing sentence discloses, and, where the source can be retrieved, whether certainty, scope, or causal strength has increased in the retelling. Compression is demanded by the form, so the aim is a reporting convention rather than a charge of carelessness.
Draft available on request.
When the evidence genuinely underdetermines a decision, more than one conclusion can be rationally held — but not any conclusion. A middle position between "one right answer" and "anything goes," separating the cases where surgeons should converge from those where reasonable people are entitled to differ.
Draft available on request.
Bayesian talk is everywhere in clinical reasoning, but the prior itself is rarely examined. Is it a summary of past frequencies, a personal degree of belief, a commitment answerable to argument — or some shifting mix? Why getting the answer wrong leads surgeons to treat contestable starting points as settled facts.
Draft available on request.
« Le mieux est l’ennemi du bien. » — Voltaire
When is one more maneuver worth it? An expected-utility exploration of the stable but imperfect proximal humerus construct. One more attempt has positive expected value only when
— but none of these quantities can be read off the fluoroscopy image; each is set by surgeon-specific priors. Defensible priors generate a band of rational stopping points, not a single answer — and the final radiograph records where a surgeon stopped, not whether stopping there was wise.
Draft available on request.
A census of how the orthopaedic literature handles the assumptions it starts from. Priors are necessarily subjective, yet the subjective Bayesian view — in the tradition of de Finetti — is not the one taken by the overwhelming majority of papers currently being written. Meanwhile "Bayes" and "Bayesian" as search terms have been increasing steadily in frequency, especially over the past five years: the vocabulary is spreading faster than the philosophical commitments it carries.
LLMs are usually benchmarked against expert humans — right or wrong relative to what a good clinician would say. This asks whether that’s the wrong yardstick, testing whether AI tracks the strength and grade of clinical recommendations rather than merely their conclusions.
High-frequency, single-item functional check-ins by text message build continuous recovery curves for individual patients — improvement velocity, total disability burden, and the point at which recovery plateaus. A move from asking whether the bone united to whether the patient recovered.
A model of pain perception and aging, led by other researchers; my role is as a collaborator.
An annotated catalogue of exemplar papers in ten themes, each with a short analysis and my own opinion set apart from it.