Skip to main navigation Skip to search Skip to main content

Performance of Open-Source Large Language Models to Extract Symptoms from Clinical Notes

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

In this study, we examined how well the open-source foundational large language models (LLMs) can extract symptoms and signs (S&S), along with their corresponding ICD-10 codes, from clinical notes found in the public MTSamples dataset. The dataset comprising notes of patients with genitourinary conditions was manually annotated to compare the S&S extraction results with outputs generated by LLMs. We assessed three versions of the Llama model—Llama 3.1-13B, Llama 3.3-70B, and Me-Llama-13B—focusing on their consistency, runtime, and performance. Each model was tested on two tasks: (1) S&S extraction and (2) ICD-10 code generation. Our findings indicate that Llama 3.3-70B performed the best overall. With fast runtime and high consistency, it achieved an average recall of 0.87 and an average precision of 0.71 for S&S extraction, as well as an average recall of 0.71 and an average precision of 0.54 for ICD-10 code generation.

Original languageEnglish (US)
Title of host publicationMEDINFO 2025 - Healthcare Smart x Medicine Deep
Subtitle of host publicationProceedings of the 20th World Congress on Medical and Health Informatics
EditorsMowafa S. Househ, Mowafa S. Househ, Zain Ul Abideen Tariq, Mahmood Al-Zubaidi, Uzair Shah, Elaine Huesing
PublisherIOS Press BV
Pages663-667
Number of pages5
ISBN (Electronic)9781643686080
DOIs
StatePublished - Aug 7 2025
Externally publishedYes
Event20th World Congress on Medical and Health Informatics, MEDINFO 2025 - Taipei, Taiwan, Province of China
Duration: Aug 9 2025Aug 13 2025

Publication series

NameStudies in Health Technology and Informatics
Volume329
ISSN (Print)0926-9630
ISSN (Electronic)1879-8365

Conference

Conference20th World Congress on Medical and Health Informatics, MEDINFO 2025
Country/TerritoryTaiwan, Province of China
CityTaipei
Period8/9/258/13/25

Keywords

  • Large Language Models
  • Llama Models
  • Natural Language Processing
  • Symptom Extraction

ASJC Scopus subject areas

  • Biomedical Engineering
  • Health Informatics
  • Health Information Management

Fingerprint

Dive into the research topics of 'Performance of Open-Source Large Language Models to Extract Symptoms from Clinical Notes'. Together they form a unique fingerprint.

Cite this