Study tests LLMs against human readers in calf carpal radiography

Bottom line

A new retrospective reader study in Veterinary Radiology & Ultrasound tested whether multimodal large language models could help score radiographs from calves with presumptive septic carpal arthritis. The study included 50 calves ages 0 to 3 months, with one carpal radiograph per case, and compared three models, ChatGPT-5, Gemini-2.5 Pro, and Claude Sonnet-4, plus two novice veterinary surgeons against an expert consensus reference using a structured ordinal scoring framework. ChatGPT-5 was the top-performing model, with moderate agreement to the expert standard, but the strongest novice reader still outperformed all of the models overall. The authors concluded that current multimodal LLMs may have limited adjunctive value for structured review or training, but they aren't ready to replace human interpretation. (lifescience.net)

Why it matters: For veterinary professionals, the study adds a useful reality check as AI tools move into imaging workflows. Septic arthritis in calves is a clinically important condition that needs prompt recognition because delayed treatment can lead to irreversible cartilage and bone damage. The paper suggests that, at least for this use case, general-purpose multimodal LLMs may be better suited to education or secondary review than frontline diagnostic decision-making, especially when compared with even relatively inexperienced human readers using a structured rubric. (merckvetmanual.com)

What to watch: Expect follow-up work on whether veterinary-specific training, larger image sets, or narrower use cases can improve AI performance enough to make these tools more clinically dependable. (onlinelibrary.wiley.com)

Key facts

Study type
Retrospective reader study
Journal
Veterinary Radiology & Ultrasound
Condition
Presumptive septic carpal arthritis in calves
Sample size
50 calves
Age range
0 to 3 months
Models tested
ChatGPT-5, Gemini-2.5 Pro, and Claude Sonnet-4
Reference standard
Expert consensus
Top model
ChatGPT-5
Main finding
The strongest novice veterinary surgeon outperformed all models overall

A new study in Veterinary Radiology & Ultrasound asks a practical question for veterinary imaging teams: can multimodal large language models meaningfully assist with radiographic scoring in a real bovine orthopedic problem? In this case, the answer was only partly yes. In a retrospective study of presumptive septic carpal arthritis in calves, ChatGPT-5 was the best-performing model tested, but it still trailed the strongest novice human reader and showed only moderate agreement with expert consensus. Claude Sonnet-4 and Gemini-2.5 Pro performed worse, with slight overall agreement. (lifescience.net)

That matters because septic arthritis remains a high-stakes diagnosis in young calves. Earlier work on 64 calves with septic arthritis found that carpal joints were the most frequently affected, particularly in neonatal calves, and standard references emphasize that septic arthritis requires prompt treatment to reduce permanent cartilage and subchondral bone damage. More recent work also suggests that radiographic bone involvement may not be as strongly tied to poor outcome as previously assumed, depending on treatment intensity, which raises the value of consistent imaging assessment even further. (pubmed.ncbi.nlm.nih.gov)

In the new study, investigators evaluated 50 calves ages 0 to 3 months with clinical findings consistent with presumptive septic carpal arthritis. One carpal radiograph per case was scored independently by two novice veterinary surgeons and three multimodal LLMs using a Constant-based ordinal radiographic scoring framework, with expert consensus used as the operational reference standard. Novice reader 1 had the strongest overall concordance, with 55.6% exact agreement, 89.8% agreement within one score category, and a mean weighted kappa of 0.68. ChatGPT-5 came closest among the models, with 53.0% exact agreement, 82.6% agreement within one category, and a mean weighted kappa of 0.58. The authors said that performance suggests limited adjunctive value, not readiness for replacement-level use. (lifescience.net)

The findings also fit the broader pattern in veterinary imaging AI: promise, but with important caveats. Prior literature in Veterinary Radiology & Ultrasound has framed AI evaluation in veterinary radiology as a validation problem, not just a novelty exercise, and other imaging studies have shown that algorithm performance can vary substantially depending on the task, dataset, and reference standard. In other words, this calf arthritis paper doesn't argue that AI has no place in radiology. It argues that general-purpose multimodal models still need careful benchmarking before they're trusted in structured clinical interpretation. (onlinelibrary.wiley.com)

Direct outside commentary on this specific paper appears limited so far, which isn't unusual for a newly indexed specialty imaging study. Still, the message aligns with existing professional caution around digital imaging workflows. The ACVR-ECVDI teleradiology consensus statement doesn't address AI directly, but it does stress image quality, sufficient clinical information, and appropriate interpretation processes, all of which are relevant when practices consider layering AI tools into imaging review. That's an inference from the consensus guidance rather than an explicit endorsement or critique of LLM-based image scoring. (onlinelibrary.wiley.com)

Why it matters: For veterinary professionals, especially those in food animal practice, teaching hospitals, and mixed practices that may rely on less experienced readers after hours, the study suggests today's multimodal LLMs are not a substitute for trained human judgment in bovine musculoskeletal radiology. Their best near-term role may be in training, checklist support, or structured second-pass review. That's potentially useful, but it's a narrower role than some AI marketing narratives imply. For clinicians talking with pet parents or livestock clients about advanced diagnostics and prognosis, the more immediate takeaway is that careful, timely imaging interpretation still depends heavily on human expertise. (lifescience.net)

What to watch: The next questions are whether veterinary-specific model tuning, multimodal systems trained on DICOM-native datasets, or use in more constrained scoring tasks can close the gap with human readers, and whether future studies validate performance across institutions rather than in a single-center retrospective design. Until then, this paper supports a cautious, adjunctive view of LLMs in veterinary radiology, not a replacement model. (lifescience.net)

Like what you're reading?

The Feed delivers veterinary news every weekday.