An artificial intelligence (AI) model called RADAR, designed to provide broad diagnostic interpretation of abdominal CT scans, substantially outperformed existing AI systems across a wide range of diseases and clinical settings, researchers report. The findings demonstrate the potential of vision-language models to advance expert-level generalist AI tools for radiology. The goal of AI in radiology is to create an expert-level generalist system capable of providing accurate, comprehensive diagnostic support across a wide range of real-world clinical situations. This, however, remains particularly challenging for contrast-enhanced abdominal computed tomography (CT), which often involves the evaluation of dozens of organs and the identification of hundreds of potential diseases, a task that requires physicians with specialized expertise and training. Although vision-language models offer a promising alternative by learning from associations between medical images and their accompanying text, abdominal CT presents unique challenges because relevant diagnostic information is sparse and highly dependent on complex anatomical context.
Here, Qi Zhang and colleagues present Rapid Abdominal Diagnosis with AI and Radiology (RADAR) – a vision-language AI model designed to provide broad diagnostic interpretation of contrast-enhanced abdominal CT scans. Trained on a dataset of 424,911 examinations, containing 1.5 million image-text pairs and more than 15 million anatomy-specific pairs, RADAR breaks CT volumes into individual anatomical structures and links those structures to corresponding descriptions in radiology reports, using a large-scale contrastive-learning approach to improve the connection between imaging findings and diagnoses. Throughout internal and external evaluations across multiple medical centers and varied clinical scenarios, Zhang et al. found that RADAR substantially outperformed existing specialist AI models, achieving a mean AUC (the average probability that a model can correctly separate positive and negative cases across evaluations) of 0.913 across 146 abdominal CT findings, compared with 0.776 for the best competing vision-language model. The model also performed well in challenging emergency settings, despite not being specifically trained on emergency data, achieving an AUC of 0.904 across more than 27,000 emergency CT cases. According to the authors, its performance remained strong across organs, diseases, and emergency cases, including uncommon findings. Moreover, in testing in cohorts at eight external centers, RADAR maintained high accuracy (AUC 0.895), demonstrating robust generalization across diverse patients, clinical settings, and imaging protocols.
Science
An expert-level generalist AI for abdominal CT diagnosis
17-Sep-2026