•  
  •  
 

Abstract

The identification of relevant studies is a critical and time-consuming component of conducting Diagnostic Test Accuracy reviews. Traditionally, this process relies on manual screening and evaluation of large volumes of literature. This study investigates the use of generative artificial intelligence, specifically OpenAI's ChatGPT, to assist in identifying and classifying relevant studies in Diagnostic Test Accuracy reviews. We applied ChatGPT and a traditional Support Vector Machine model across four Diagnostic Test Accuracy reviews and evaluated their performance using precision, recall, and F1 score. ChatGPT achieved an average recall of 85.18% and F1 score of 78.97%, outperforming SVM in recall and overall balance, while SVM demonstrated higher precision. These findings suggest that generative artificial intelligence offers significant advantages in early-stage screening for systematic reviews. The study also highlights the role of ChatGPT in producing structured justifications aligned with the PICO framework, enhancing transparency and decision support.

Keywords

Generative AI, Diagnostic test accuracy (DTA) reviews, Natural language processing, Explainable artificial intelligence (XAI), Evidence synthesis, Information extraction

Article Type

Article

First Page

86

Last Page

93

Publication Date

6-30-2026

Share

COinS