Analysis of Large Language Model Decision Making in Hormone Receptor-Positive/Human Epidermal Growth Factor Receptor 2-Negative Early Breast Cancer
Articolo
Data di Pubblicazione:
2026
Citazione:
Analysis of Large Language Model Decision Making in Hormone Receptor-Positive/Human Epidermal Growth Factor Receptor 2-Negative Early Breast Cancer / Buonaiuto, R., Caltavituro, A., Di Rienzo, R., Grieco, A., Mangiacotti, F.P., Longobardi, A., Cantile, V., Molinaro, V., Pagliuca, M., Buono, G., De Placido, P., Pietroluongo, E., Forestieri, V., Martinelli, C., Di Lauro, V., Leo, L., D'Aiuto, M., Bianchini, G., Criscitiello, C., Bianco, R., et al.. - In: JCO CLINICAL CANCER INFORMATICS. - ISSN 2473-4276. - 10:10(2026). [10.1200/CCI-25-00230]
Abstract:
PURPOSE: To assess the ability of GPT-4o in adjuvant treatment decision making in hormone receptor-positive (HR+)/human epidermal growth factor receptor 2-negative (HER2-) early breast cancer by comparing its recommendations with those of clinicians including Oncotype DX data, and to explore its potential as a decision-support tool in routine clinical practice. METHODS: We compared clinician and GPT-4o recommendations in patients tested with Oncotype DX in routine practice at the University of Naples Federico II (n = 607, cohort 1 [C1]) and within the prospective, multicenter PRO BONO study (n = 237, cohort 2 [C2]). Pre- and post-Oncotype DX treatment recommendations were categorized as chemotherapy (CT) + endocrine therapy (ET) or ET alone. Concordance between clinician and GPT-4o recommendations was assessed using agreement rates and Cohen's kappa. The accuracy of Oncotype DX results was evaluated using the AUC metric. RESULTS: The agreement between clinicians and GPT-4o in pretest recommendations was 68% (kappa, 0.381 [95% CI, 0.31 to 0.45], P < .001) in C1 and 70% (0.401 [95% CI, 0.29 to 0.52], P < .001) in C2. Before Oncotype DX, clinicians recommended CT more frequently than GPT-4o for C1 (58% v 38%) and C2 (53% v 43%). Post-test agreement increased to 93% (0.814 [95% CI, 0.76 to 0.87], P < .001) in C1 and 90% (0.741 [95% CI, 0.64 to 0.84], P < .001) in C2. The agreement between pre- and post-Oncotype DX treatment recommendations for clinicians was 56% and 63% versus 68% and 60% for GPT-4o in C1 and C2, respectively. GPT-4o showed higher accuracy in predicting low than high genomic risk in postmenopausal patients (87% v 43% in C1; 85% v 45% in C2, P < .001) and low versus intermediate and high risk in premenopausal patients in both cohorts (P < .001). CONCLUSION: The agreement between clinicians and GPT-4o in pretest recommendations was modest but improved post-test, highlighting the importance of multigene testing and the potential of large language models in clinical decision making.
Tipologia CRIS:
1.1 Articolo in rivista
Elenco autori:
Buonaiuto, R.; Caltavituro, A.; Di Rienzo, R.; Grieco, A.; Mangiacotti, F. P.; Longobardi, A.; Cantile, V.; Molinaro, V.; Pagliuca, M.; Buono, G.; De Placido, P.; Pietroluongo, E.; Forestieri, V.; Martinelli, C.; Di Lauro, V.; Leo, L.; D'Aiuto, M.; Bianchini, G.; Criscitiello, C.; Bianco, R.; Del Mastro, L.; De Laurentiis, M.; Arpino, G.; De Angelis, C.; Giuliano, M.
Link alla scheda completa:
Pubblicato in: