Presentation
Can AI Effectively Screen Complex Literature: A Case Study of Driver Engagement Research
DescriptionAI-assisted tools are increasingly used to support literature screening, yet their effectiveness in complex, interdisciplinary domains remains unclear. This study evaluates the performance of an AI-assisted systematic review tool (Elicit) during abstract screening for a human factors research question on driver engagement in Level 2 automation. Elicit retrieved 500 records and applied screening criteria to generate inclusion decisions and ratings. We manually reviewed 335 records to assess the AI tool’s precision, discriminative power, and uncertainty. Results showed low overall precision (10.71%). However, among the papers ultimately included through complete manual review, 94% had received the AI tool’s highest rating (4.9), suggesting that the rating may help prioritize records for subsequent human screening even though overall precision is limited. Analysis of screening criteria revealed that methodological criteria such as human participants and empirical evidence, were easily satisfied but weak at filtering irrelevant studies, whereas topic-related criteria such as automation level and engagement effectiveness exhibited stronger discriminative potential but greater uncertainty. These findings suggest that AI still struggles with conceptually complex distinctions. We conclude that AI-assisted screening can improve efficiency but still requires human expertise to resolve conceptual ambiguity, particularly in interdisciplinary domains such as human factors.
Event Type
Lecture
TimeTuesday, October 20th1:30pm - 1:50pm PDT
Location
