28/05/2026
Most AI studies in medical imaging report strong performance on their own data. The real question is whether those models work just as well in the real world.
A study published in Surgeries by Nasef, Sawiris, Girgis and Toma (New York Institute of Technology) puts this question at the center, evaluating deep learning models for knee osteoarthritis grading across two separate datasets. The findings show that without external validation, AI models can appear far more reliable than they actually are. 🔬
Full article: https://brnw.ch/21x2T1p
Background: This study evaluated the performance of machine learning models trained on two different datasets of knee X-ray images annotated with Kellgren–Lawrence grades. Methods: Learning curves indicated that one model experienced poor training, characterized by underfitting, while the other mo...