Hand (N Y). 2026 Aug 4:15589447261465836. doi: 10.1177/15589447261465836. Online ahead of print.
ABSTRACT
BACKGROUND: Accurate assessment of distal radius fracture stability is essential for appropriate triage and timely referral to hand specialists. The LaFontaine criteria provide a structured radiographic framework for predicting instability but are not routinely reported. Recent advances in multimodal large language models (LLMs) capable of direct image interpretation have generated interest in their potential role as adjunctive diagnostic tools. However, their real-world performance in structured radiographic assessment remains unclear.
METHODS: A cross-sectional diagnostic accuracy and agreement study was performed using 20 distal radius fracture radiographs. Five hand surgeons independently assessed each case for the 5 LaFontaine criteria, with majority agreement serving as the reference standard. Two publicly available multimodal LLMs (ChatGPT and Claude) were evaluated using a standardized, single-prompt approach designed to approximate real-world use. Both models were provided identical radiographs and clinical prompts then asked to classify each criterion and overall fracture stability. Agreement and diagnostic performance were calculated relative to surgeon consensus.
RESULTS: Hand surgeons demonstrated high consistency in identifying LaFontaine criteria. Agreement between LLMs and clinician consensus varied across individual features, with several criteria showing limited agreement. Both models achieved similar overall accuracy for fracture stability classification (0.75; 95% CI, 0.56-0.94). ChatGPT demonstrated moderate agreement with surgeon consensus, while Claude showed fair agreement. Agreement between clinician-determined instability and independent operative recommendations was moderate.
CONCLUSIONS: Multimodal LLMs demonstrated variable and generally limited agreement with clinician consensus in classifying fracture stability. Performance across several criteria approached chance levels, and specificity was limited. These findings should be interpreted as exploratory and hypothesis-generating; current models are not yet reliable for clinical decision-making and require substantial validation before potential use as adjunctive triage tools.
LEVEL OF EVIDENCE: Diagnostic Level III.
PMID:42552768 | DOI:10.1177/15589447261465836

