Radiology artificial intelligence: a systematic review and evaluation of methods
Key points
- This systematic review covered deep learning studies in clinical radiology published from 2015 to 2019: 11,083 search results, 767 full texts reviewed and 535 articles included.
- 98% of studies were retrospective, only 13 were prospective, and a power analysis was referenced in only 5%.
- The most common features were MRI (37%), neuroradiology (24%), a segmentation task (39%), supervised learning (88%) and the UNet architecture (14%).
- Median performance was a Dice of 0.89, an AUC of 0.903 and an accuracy of 89.4; in 77 studies allowing direct comparison, performance decreased on average by 6% at external validation (range 4% improvement to 44% reduction).
- 17% of studies offered no performance comparison, 28% offered no explainability, and the source code or model was freely available in only 14%.
Overview
The RAISE systematic review surveyed studies that apply deep learning to clinical radiology, to identify the clinical questions being asked and to evaluate the methods used. Its main finding is that many papers report expert-level results, but most apply a narrow range of techniques to a narrow selection of use cases, and the literature is dominated by retrospective cohort studies with limited external validation and a high potential for bias. The review followed PRISMA with a prospectively registered protocol (PROSPERO: CRD42020154790) and searched MEDLINE and EMBASE from 1 January 2015 to 31 December 2019.
Use cases and modalities
The number of included studies rose every year: 14, 31, 84, 170 and 237 from 2015 to 2019. Cancer imaging was the most common use case (156 studies, 29%), with a further 45 (8%) specifically examining pulmonary nodules. Neuroradiology (127, 24%) and chest (92, 17%) were the leading subspecialties, and MR (200, 37%) and CT (155, 29%) the leading modalities. The main tasks were segmentation (211, 39%), classification (171, 31%) and identification (74), and no study had outcomes data.
Study design and ground truth
525 studies (98%) were retrospective, 13 were prospective and only 1 used a prospective real-world evaluation. Ground truth came from the radiologic report in 66%, from an expert panel or other specialty reports in 19% and from pathology in 9%. Performance was compared with a state-of-the-art model in 37% and with radiologists in 31%, while 17% had no comparison.
Data and methods
The median number of patients was 460 (range 3 to 313,318) and the median number of images was 1993. Data came from one hospital in 44% and from a public dataset in 32%, and external validation was used in 169 studies (31%). Transfer learning was used in 46%, and U-Net and custom architectures showed very similar Dice performance (0.877 and 0.895).
Performance, explainability and open access
In 60 of 77 studies (78%), the drop at external validation was 10% or less; together with average metrics at roughly 90% of maximum, this raises the possibility of publication bias. Explainability was most often provided as cases or examples (49%), was absent in 28%, used visualisations and saliency maps in 18%, and never used counterfactual examples. Code or model was available in 14% and data in 38%. The authors call for prospective registration, reporting guidelines such as those from CONSORT and the RSNA, external validation and higher-level studies with outcomes data.
Frequently asked questions
How well do radiology AI models perform outside the institution that developed them?
In the 77 studies that allowed a direct comparison, performance decreased on average by 6% at external validation, ranging from a 4% improvement to a 44% reduction.
What type of study dominated the literature?
Retrospective cohort studies: 98% of included studies were retrospective, and only 13 were prospective.
What do the authors recommend?
Prospective study registration, reporting guidelines, more external validation and higher-level studies with outcomes data to show that the claims translate into benefits for patients.
Source
Kelly BS, Judge C, Bollard SM, Clifford SM, Healy GM, Aziz A, et al. Radiology artificial intelligence: a systematic review and evaluation of methods (RAISE). European Radiology 2022;32(11):7998-8007. DOI: 10.1007/s00330-022-08784-6. Open access under a Creative Commons Attribution 4.0 licence; a correction was published later. This page is a summary prepared by Medpresso from the original publication and is not a substitute for the full text or for medical advice.