1. Dippel and colleagues evaluated AI detection models’ ability to identify infrequent pathological changes on gastrointestinal (GI) biopsies.
2. The best-performing model detected unfamiliar malignancies with high accuracy but missed some subtle histological changes.
Evidence Rating Level: 2 (Good)
Study Rundown: Diagnostic pathology faces serious challenges due to both a shortage of qualified pathologists and an increasing cancer diagnosis burden. AI models have demonstrated the potential to improve cancer detection by analyzing biopsies. Dippel and colleagues evaluated whether AI models could identify less common histopathological findings without specific training on the target diseases. Researchers trained AI models using two large real-world datasets of GI biopsies, with the 10 most common findings accounting for 90% of the cases. The models’ performances were assessed by comparing them to clinical diagnoses and pathologist annotations. The study found that the best-performing model achieved areas under the receiver operating characteristic curve (AUROC) of 0.950 for gastric biopsies and 0.910 for colonic biopsies, with particularly strong discrimination for malignancies. Additionally, the models generated heatmaps to highlight anomalous areas for pathologist review. However, subtle colonic changes remained difficult to detect. This study demonstrated that AI has the potential to support pathologists by flagging anomalous cases and reducing missed diagnoses.
Click here to read the study in NEJM AI
Relevant Reading: Artificial intelligence applications in histopathology
In-Depth [retrospective cohort]: 5,423 tissue slides, yielding approximately 17 million image patches, were included in the study. Clinical reports associated with the slides included pathologist-verified annotations. These annotations delineated regions of interest containing the actual diagnosis-defining anomalies. Models were trained to distinguish common target-organ patterns from other tissues, and low-grade adenomas were excluded from training to improve detection of high-grade changes. Performance metrics included AUROC for gastric and colonic biopsies. Outlier exposure achieved slide-level AUROCs of 95.04% (standard deviation ± 0.54) for stomach and 91.01% (± 0.69) for colon. Corresponding patch-level AUROCs were 91.37% (± 0.34) and 90.47% (± 0.33). For malignancies, slide-level AUROCs reached 97.72% (± 0.44) and 96.97% (± 0.61), respectively. Without retraining, external slide-level AUROCs were 94.50% for stomach and 85.88% for colon. However, the models faced challenges in reliably recognizing pseudomelanosis coli and intestinal spirochetosis. This study was limited by the difficulty of detecting subtle or architectural changes with small image patches, validation at only two hospitals, and exclusion of esophageal, small-intestinal, and anal biopsies. No prospective evaluation established reductions in missed diagnoses or pathologist workload. These results support further evaluation as a tool for flagging cases for expert review, rather than autonomous diagnosis of individual rare diseases. Overall, this study demonstrated that while AI cannot autonomously diagnose individual rare diseases, it has the potential to be a useful tool for flagging cases for expert review.
Image: PD
©2026 2 Minute Medicine, Inc. All rights reserved. No works may be reproduced without expressed written consent from 2 Minute Medicine, Inc. Inquire about licensing here. No article should be construed as medical advice and is not intended as such by the authors or by 2 Minute Medicine, Inc.




