Study flags undertriage risk for LLMs in feline fracture triage

Bottom line

Large language models showed uneven performance when asked to decide whether cats with metacarpal and metatarsal fractures needed surgery, according to a new JAVMA study of 73 retrospective cases reviewed against a reference standard set by two board-certified veterinary orthopedic surgeons. The researchers evaluated five LLMs using the same case materials and found that performance varied by model, with a consistent safety concern: a tendency to undertriage some cases that experts judged should receive surgical management. The paper adds to a growing veterinary AI literature suggesting these tools can appear capable in structured decision tasks while still missing clinically important cases. (pubmed.ncbi.nlm.nih.gov)

Why it matters: For veterinary professionals, the key takeaway isn’t just that accuracy differed across models, but that the error pattern leaned toward false reassurance. In fracture triage, undercalling a surgical case could delay referral, fixation, or more intensive monitoring, especially in injuries that may already be easy to miss on initial exam. That makes this less a story about whether LLMs can help and more a story about where guardrails are needed if clinics use them for decision support, client communication, or workflow triage. Prior veterinary and human-health AI studies have reached a similar caution: aggregate performance can mask subgroup failures, and external validation is essential before real-world deployment. (pubmed.ncbi.nlm.nih.gov)

What to watch: Expect follow-up work on prospective validation, prompt design, and whether specialty-tuned or multimodal models can reduce undertriage without creating new safety risks. (pubmed.ncbi.nlm.nih.gov)

Key facts

Study type
JAVMA study
Species
Cats
Condition
Metacarpal and metatarsal fractures
Clinical question
Whether surgery was needed
Cases reviewed
73 retrospective cases
Reference standard
Expert consensus from two board-certified veterinary orthopedic surgeons
Models tested
Five large language models
Main finding
Performance varied by model and some cases were undertriaged
Safety concern
False reassurance for cases judged to need surgical management

A new JAVMA study raises a practical warning for clinics experimenting with generative AI: large language models did not perform consistently when triaging feline metacarpal and metatarsal fractures for surgery, and they showed a systematic tendency to undertriage some cases that veterinary orthopedic specialists believed needed operative management. The study reviewed 73 retrospective cases collected between December 2023 and February 2025 and compared five LLMs against an expert consensus reference standard from two board-certified veterinary orthopedic surgeons. (pubmed.ncbi.nlm.nih.gov)

That matters because these distal limb fractures can be deceptively subtle. Background orthopedic references note that non-displaced feline metacarpal and metatarsal fractures may be easy to miss, and prior feline fracture literature has shown that treatment decisions can hinge on fracture configuration, the number of bones involved, displacement, and soft-tissue injury. Surgical stabilization can produce strong functional outcomes in appropriately selected cases, but conservative management also has a role in some patients, making triage judgment especially important. (sciencedirect.com)

The new paper fits into a fast-expanding body of work from the same research network and from the broader veterinary AI field. Recent studies have tested LLMs in veterinary anesthetic planning, ocular diagnosis, and emergency triage, often finding that model performance is highly task-dependent rather than uniformly reliable. A separate Veterinary Record study on emergency triage found that AI models were reasonably sensitive for severe emergencies but still misclassified many lower-acuity cases, while another 2026 trauma-triage study in human medicine warned that strong overall scores can conceal clinically meaningful undertriage in specific subgroups. (pubmed.ncbi.nlm.nih.gov)

What stands out here is the direction of the error. In many triage settings, overtriage creates inefficiency; undertriage creates delay and can be harder to catch. For a cat that actually needs surgical stabilization, a conservative recommendation from an LLM could influence referral timing, client counseling, analgesia planning, or expectations around splinting and follow-up. That’s especially relevant in first-opinion settings, where a tool that sounds confident may be used to support decisions before specialist review. More broadly, external validation work on commercial veterinary radiology AI has already shown that promising systems can perform poorly on real-world general practice cases, reinforcing the need for caution before deployment. (pubmed.ncbi.nlm.nih.gov)

Direct expert reaction to this specific study was limited in publicly available sources, but the broader commentary in veterinary AI has become more consistent: these systems may have value as supervised adjuncts, not autonomous clinical decision-makers. A JAVMA framework paper on safe AI deployment in veterinary medicine argues for governance, validation, and workflow safeguards, and recent ethics commentary has similarly emphasized that decision support should not displace professional judgment in situations where patient safety depends on nuanced interpretation. (pubmed.ncbi.nlm.nih.gov)

Why it matters: For veterinary teams, this study is a reminder to evaluate AI tools by failure mode, not just average accuracy. A model that performs “well” overall may still be unsafe if its mistakes consistently downplay cases that need escalation. In orthopedic triage, that means clinics should be wary of using general-purpose LLMs to decide when a fracture can be managed conservatively versus when it should move quickly to surgical consultation. If LLMs are used at all, the safer near-term role may be administrative summarization, draft client education, or flagged second looks, with clear human oversight. (pubmed.ncbi.nlm.nih.gov)

What to watch: The next step will likely be prospective and multicenter validation, including testing whether structured prompts, image-enabled models, or specialty-specific fine-tuning improve reliability. Just as important will be reporting on calibration, repeatability, and subgroup errors, because in triage work, the most consequential question isn’t whether AI is sometimes right, but whether clinicians can predict when it will be wrong. (pubmed.ncbi.nlm.nih.gov)

Like what you're reading?

The Feed delivers veterinary news every weekday.