Evaluating Clinical Symptom Extraction and Reasoning in Local Large Language Models: A Proof-of-Concept Study of Symptom-Specific Performance in Cardiology
Shun Kitamura, Eisuke Amiya, Yoshihiro Izawa, Risa Kishikawa, Takenobu Shimada, Junichi Ishida, Satoshi Kodera, Norihiko Takeda
Source record
Source: Crossref
Published: Sep 17, 2026
DOI: 10.21203/rs.3.rs-10844001/v1
Open original source ↗Source abstract
Abstract Background Patients with heart failure present with various symptoms, and comprehensively identifying and accurately documenting them in clinical registries or structured databases requires substantial manual effort. Locally deployed large language models (LLMs), which process data entirely within institutional infrastructure, have enabled automated extraction of structured information from unstructured clinical text. However, symptom-specific extraction performance and reasoning-trace analyses of the underlying decision-making process remain largely unexplored. We utilized multiple locally deployed LLMs for symptom extraction from cardiology discharge summaries, examining how extraction performance varied across individual symptoms and characterizing the inferential processes through reasoning traces. Methods Ten Japanese-language clinical summaries of patients with advanced heart failure evaluated for heart transplantation listing were analyzed, yielding 80 reference-standard symptom-case labels (10 cases × 8 symptoms). Gemma3-27b and Qwen3.5-9b were run locally on a single GPU to extract the presence or absence of eight symptoms required for transplant listing in Japan. Two prompting conditions were evaluated: (1) extraction of all symptoms from the full summary, and (2) an additional instruction to disregard symptoms unrelated to the index hospitalization. Phase 1 evaluated four non-reasoning conditions, and phase 2 added chain-of-thought reasoning to each condition. Results are reported descriptively, without formal statistical testing. Results Overall accuracy across the four non-reasoning conditions ranged from 87.5% to 91.3%, with Qwen3.5 showing greater sensitivity to prompt modification than Gemma3. In symptom-specific evaluation across all conditions, palpitations and fatigability showed the highest false-positive rates (25% and 38%, respectively). Review of the reasoning outputs showed that, in some cases, both models inferred palpitations from electrocardiographic findings or documented tachycardia, and fatigability from the underlying diagnosis of heart failure, despite the absence of explicit symptom documentation. Conclusions Symptom extraction performance varies according to the linguistic and clinical characteristics of individual symptoms. Reasoning-trace analysis further showed that LLMs could infer symptom presence from objective clinical findings rather than relying solely on explicitly documented patient complaints.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.