A large language model for risk-of-bias assessment in systematic reviews of prognosis studies in clinical neurology
W skrócie
[Preprint - wstępne wyniki] Naukowcy sprawdzili, czy sztuczna inteligencja może automatycznie oceniać jakość badań naukowych dotyczących rokownika chorób neurologicznych (epilepsji, udarów, urazów mózgu). Chociaż zgoda między oceną AI a oceną lekarzy była umiarkowana, wyniki sugerują, że automatyczna ocena mogłaby być równie dokładna co ocena przez różnych lekarzy. Autorzy uważają, że po ulepszczeniach metodologicznych ta technologia mogłaby zaoszczędzić czas i pieniądze przy tworzeniu przeglądu badań naukowych w neurologii.
Oryginalny abstract (angielski)
Background: Risk-of-bias (ROB) assessments represent an integral component of systematic reviews. However, this task is often highly repetitive, time-consuming, and may lack inter-rater consistency. Large language models (LLMs) offer opportunities for automation in systematic reviews, which may expedite and enhance the quality and consistency of research synthesis. Methods: Using zero-shot prompting, we designed an LLM-based pipeline as a virtual mimic of a human reviewer for the Quality in Prognosis Studies (QUIPS) framework. Then, focusing on prognostic research in a single discipline (neurology), we applied this pipeline to articles included in previously published systematic reviews. We studied inter-rater agreement between both (1) the LLM and the original human ROB assessments and (2) between original human ROB assessments. Results: 298 articles from 15 reviews across three domains (epilepsy, traumatic brain injury, stroke) were included. We demonstrate the feasibility of a tailored, prompt-engineered LLM pipeline for automating ROB assessments with the QUIPS tool. While LLM-human agreement was limited (Cohen's weighted kappa; = 0.22, 95% CI, 0.12 - 0.33), our data tentatively suggest, based on a small sample (n=5), that it may not be inferior to human-human agreement (Cohen's weighted kappa; = -0.25, 95% CI, -1.04 - 0.54). Wilcoxon signed-rank tests were statistically significant (p < 0.05) across four bias domains and for the overall risk scores, and rank-biserial correlations demonstrated human raters' tendency to assign higher risk scores than LLM counterparts. Conclusions: With targeted methodological refinements - including standardization of QUIPS implementation and validation against expert ratings - automated ROB assessments may meaningfully reduce time and cost of systematic reviews of prognosis studies in neurology and beyond.