Information extraction from medical case reports using OpenAI InstructGPT

Sciannameo, Veronica; Jahier Pagliari, Daniele; Urru, Sara; Grimaldi, Piercesare; Ocagli, Honoria; Ahsani-Nasab, Sara; Rosanna Irene Comoretto,; Gregori, Dario; Berchialla, Paola

doi:10.1016/j.cmpb.2024.108326

Background and objective: Researchers commonly use automated solutions such as Natural Language Processing (NLP) systems to extract clinical information from large volumes of unstructured data. However, clinical text's poor semantic structure and domain-specific vocabulary can make it challenging to develop a one-size-fits-all solution. Large Language Models (LLMs), such as OpenAI's Generative Pre-Trained Transformer 3 (GPT-3), offer a promising solution for capturing and standardizing unstructured clinical information. This study evaluated the performance of InstructGPT, a family of models derived from LLM GPT-3, to extract relevant patient information from medical case reports and discussed the advantages and disadvantages of LLMs versus dedicated NLP methods. Methods: In this paper, 208 articles related to case reports of foreign body injuries in children were identified by searching PubMed, Scopus, and Web of Science. A reviewer manually extracted information on sex, age, the object that caused the injury, and the injured body part for each patient to build a gold standard to compare the performance of InstructGPT. Results: InstructGPT achieved high accuracy in classifying the sex, age, object and body part involved in the injury, with 94%, 82%, 94% and 89%, respectively. When excluding articles for which InstructGPT could not retrieve any information, the accuracy for determining the child's sex and age improved to 97%, and the accuracy for identifying the injured body part improved to 93%. InstructGPT was also able to extract information from non-English language articles. Conclusions: The study highlights that LLMs have the potential to eliminate the necessity for task-specific training (zero-shot extraction), allowing the retrieval of clinical information from unstructured natural language text, particularly from published scientific literature like case reports, by directly utilizing the PDF file of the article without any pre-processing and without requiring any technical expertise in NLP or Machine Learning. The diverse nature of the corpus, which includes articles written in languages other than English, some of which contain a wide range of clinical details while others lack information, adds to the strength of the study.

Information extraction from medical case reports using OpenAI InstructGPT / Sciannameo, V., JAHIER PAGLIARI, D., Urru, S., Grimaldi, P., Ocagli, H., Ahsani-Nasab, S., Irene Comoretto, R., Gregori, D., Berchialla, P.. - In: COMPUTER METHODS AND PROGRAMS IN BIOMEDICINE. - ISSN 1872-7565. - ELETTRONICO. - 255:(2024). [10.1016/j.cmpb.2024.108326]

Information extraction from medical case reports using OpenAI InstructGPT

Veronica Sciannameo;Daniele Jahier Pagliari;Sara Urru;Piercesare Grimaldi;Honoria Ocagli;Sara Ahsani-Nasab;Rosanna Irene Comoretto;Dario Gregori;Paola Berchialla

2024

Abstract

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno del prodotto
	
				2024
			
	Codice DOI
	
				https://dx.doi.org/10.1016/j.cmpb.2024.108326
			
	Titolo della Rivista
	
				COMPUTER METHODS AND PROGRAMS IN BIOMEDICINE
			
	Appare nelle tipologie
	
				1.1 Articolo in rivista

File in questo prodotto:

File	Dimensione	Formato
1-s2.0-S0169260724003195-main.pdf accesso aperto Descrizione: Versione Editoriale Tipologia: 2a Post-print versione editoriale / Version of Record Licenza: Creative commons Dimensione 345.87 kB Formato Adobe PDF Visualizza/Apri	345.87 kB	Adobe PDF	Visualizza/Apri

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/2991607

PORTO @ Archivio Istituzionale della Ricerca

Information extraction from medical case reports using OpenAI InstructGPT

Veronica Sciannameo;Daniele Jahier Pagliari;Sara Urru;Piercesare Grimaldi;Honoria Ocagli;Sara Ahsani-Nasab;Rosanna Irene Comoretto;Dario Gregori;Paola Berchialla

2024

Abstract

Scheda breve Scheda completa Scheda completa (DC)

Pubblicazioni consigliate

Informazioni

Conferma cancellazione

Scheda breve

Scheda completa

Scheda completa (DC)