Guidelines, Consensus Statements, and Standards for the Use of Artificial Intelligence in Medicine
Key points
- The review assessed the quality of guidelines, consensus statements and standards on artificial intelligence (AI) in medicine to provide a foundation for future AI guideline development.
- It included 19 guideline articles, 14 consensus statements and 3 standards published between 2019 and 2022, found in 7 databases searched up to 6 April 2022.
- Methodological quality was assessed with AGREE II: the mean overall score was 4.0 on a 7-point scale (range 2.2-5.5).
- Reporting quality was assessed with RIGHT: the mean overall reporting rate was 49.4% (range 25.7%-77.1%).
- Quality differed considerably between documents, and the authors made recommendations to improve methodological and reporting quality.
Overview
AI is increasingly used in health care, and many guidelines, consensus statements and standards have been produced on its use. This systematic review, registered in PROSPERO (CRD42022321360), appraised their methodological and reporting quality and compared their content. This overview is based on the abstract and selected results of the review.
What the documents covered
The included documents addressed disease screening, diagnosis and treatment, reporting of AI intervention trials, AI imaging development and collaboration, AI data application, and AI ethics governance and applications. Examples include AI screening for retinopathy, AI in oesophageal cancer diagnosis and treatment, 3D visualisation of lung nodules for localisation and surgical planning, evaluation of commercial AI imaging solutions, building ophthalmology image databases and constructing medical data sets.
Quality findings
The overall aims, health goals and expected outcomes were clearly stated in each document. All but 5 articles fully described how recommendations were formed, and 14 articles explicitly linked recommendations to supporting evidence.
FAQ
How good are current AI guidelines in medicine?
Quality varies widely: the mean AGREE II overall score was 4.0 out of 7 (range 2.2-5.5) and the mean RIGHT reporting rate was 49.4%.
Which tools were used to assess quality?
AGREE II for methodological quality and the RIGHT checklist (7 domains, 22 items and 35 subitems) for reporting quality.
Source
Wang Y, Li N, Chen L, et al. Guidelines, Consensus Statements, and Standards for the Use of Artificial Intelligence in Medicine: Systematic Review. J Med Internet Res 2023;25:e46089. DOI: 10.2196/46089. Open access under CC BY 4.0. Summary prepared by Medpresso from the original publication; it is not a substitute for the full text or for medical advice.