Abstract
The identification of relevant studies is a critical and time-consuming component of conducting Diagnostic Test Accuracy reviews. Traditionally, this process relies on manual screening and evaluation of large volumes of literature. This study investigates the use of generative artificial intelligence, specifically OpenAI's ChatGPT, to assist in identifying and classifying relevant studies in Diagnostic Test Accuracy reviews. We applied ChatGPT and a traditional Support Vector Machine model across four Diagnostic Test Accuracy reviews and evaluated their performance using precision, recall, and F1 score. ChatGPT achieved an average recall of 85.18% and F1 score of 78.97%, outperforming SVM in recall and overall balance, while SVM demonstrated higher precision. These findings suggest that generative artificial intelligence offers significant advantages in early-stage screening for systematic reviews. The study also highlights the role of ChatGPT in producing structured justifications aligned with the PICO framework, enhancing transparency and decision support.
Keywords
Generative AI, Diagnostic test accuracy (DTA) reviews, Natural language processing, Explainable artificial intelligence (XAI), Evidence synthesis, Information extraction
Article Type
Article
First Page
86
Last Page
93
Publication Date
6-30-2026
Recommended Citation
Alharbi, Amal Hamed
(2026)
"Exploring the Capabilities of Generative AI in Supporting Diagnostic Test Accuracy Reviews,"
Journal of King Abdulaziz University: Computing and Information Technology Sciences: Vol. 15:
Iss.
1, Article 6.
DOI: https://doi.org/10.64064/1658-6336.1024
