Start Over

Advancing Italian Biomedical Information Extraction with Large Language Models: Methodological Insights and Multicenter Practical Application

Authors :: Crema, Claudio
Buonocore, Tommaso Mario
Fostinelli, Silvia
Parimbelli, Enea
Verde, Federico
Fundarò, Cira
Manera, Marina
Ramusino, Matteo Cotta
Capelli, Marco
Costa, Alfredo
Binetti, Giuliano
Bellazzi, Riccardo
Redolfi, Alberto
Publication Year :: 2023
Publisher :: arXiv, 2023.
Abstract: The introduction of computerized medical records in hospitals has reduced burdensome operations like manual writing and information fetching. However, the data contained in medical records are still far underutilized, primarily because extracting them from unstructured textual medical records takes time and effort. Information Extraction, a subfield of Natural Language Processing, can help clinical practitioners overcome this limitation, using automated text-mining pipelines. In this work, we created the first Italian neuropsychiatric Named Entity Recognition dataset, PsyNIT, and used it to develop a Large Language Model for this task. Moreover, we conducted several experiments with three external independent datasets to implement an effective multicenter model, with overall F1-score 84.77%, Precision 83.16%, Recall 86.44%. The lessons learned are: (i) the crucial role of a consistent annotation process and (ii) a fine-tuning strategy that combines classical methods with a "few-shot" approach. This allowed us to establish methodological guidelines that pave the way for future implementations in this field and allow Italian hospitals to tap into important research opportunities.

Subjects :: FOS: Computer and information sciences
Computer Science - Machine Learning
Computer Science - Computation and Language
J.3
Artificial Intelligence (cs.AI)
Computer Science - Artificial Intelligence
I.2.7
Computation and Language (cs.CL)
Machine Learning (cs.LG)

Details

Database :: OpenAIRE
Accession number :: edsair.doi.dedup.....37fcacf4a83a9abc62614e6fea403290
Full Text :: https://doi.org/10.48550/arxiv.2306.05323

Full Text Access

View/download PDF

Tools

Email
Cite

Printer

Authors Abstract Subjects Details

Searchworks

Select search scope, currently: Articles

Catalog

books, media & more in Jio Institute collections

Articles

journal articles & other e-resources

Advancing Italian Biomedical Information Extraction with Large Language Models: Methodological Insights and Multicenter Practical Application

Abstract

Subjects

Details

Tools

Searchworks

Select search scope, currently: Articles Catalog books, media & more in Jio Institute collections Articles journal articles & other e-resources

Advancing Italian Biomedical Information Extraction with Large Language Models: Methodological Insights and Multicenter Practical Application

Abstract

Subjects

Details

Tools

Select search scope, currently: Articles

Catalog

books, media & more in Jio Institute collections

Articles

journal articles & other e-resources