Overview
Medical large language models (LLMs) are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records). Yet how these data types differentially shape model capabilities has remained unclear. A new study, accepted for oral presentation at NLPCC 2026, addresses this question through token-matched experiments.
Methodology
The researchers varied the didactic-to-clinical ratio in training data and analyzed how data composition affects performance, capability profiles, and error patterns across both knowledge-intensive and clinic-oriented tasks.
Key Findings
Asymmetric Transfer Across Task Types
- Clinical data improves clinic-oriented tasks while remaining competitive on knowledge-intensive ones.
- Didactic data mainly improves knowledge-intensive tasks.
The Knowing-Doing Gap
Error analysis suggests a knowing-doing gap, where improvements in knowledge recall do not reliably generalize to clinical reasoning.
Data Efficiency and Optimal Ratios
- Modest amounts of clinical data yield most of the gains on EHR-grounded tasks.
- The optimal mixture ratio varies with the knowledge and clinical reasoning demands of downstream tasks.
Implications
These findings suggest that medical LLM data curation should be application-driven, with higher proportions of clinical data preferred for reasoning-intensive use cases.
Paper Details
- arXiv ID: arXiv:2609.22161 [cs.AI]
- Submitted: 26 Aug 2026
- Authors: Yuzheng Fan, Haochun Wang, Sendong Zhao, Xiao Han, Ming Ma, Bing Qin
- Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
- Comments: Accepted by NLPCC 2026 oral
- DOI: https://doi.org/10.48550/arXiv.2609.22161
via ArXiv AI
