ChestPheNoT: Deployable, Auditable Label-Status-Evidence Extraction from Radiology Reports
Authors: Kai Yu, Chenyu Zhu, Zaifu Zhan, Meijia Song, Min Zeng, Xiaoyi Chen, Mingquan Lin, Rui Zhang
Venue: Accepted at IEEE Healthcom 2026
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
arXiv: arXiv:2609.31629 [cs.CL]
Code: https://github.com/yukkai/ChestPheNoT
Overview
Structured phenotype extraction from radiology reports underpins cohort construction, quality auditing, and clinical analytics. However, practical deployment imposes two hard constraints: inference must run locally, and predictions must be auditable. Expert annotations, meanwhile, remain scarce.
Existing approaches fall short of these requirements. Conventional labelers output structured findings and assertion states but provide no supporting evidence. API-hosted large language models, by contrast, may be unusable when clinical text is prohibited from leaving institutional infrastructure.
Introducing ChestPheNoT
We present CHESTPHENOT, a compact 0.5–3B parameter language model that jointly extracts:
- Finding labels — the phenotype of interest
- Three-class status — present, absent, or uncertain
- Verbatim supporting evidence spans — the exact source text justifying each prediction
This joint formulation makes every prediction directly traceable to the report text, enabling auditing without a separate explanation step.
Training Recipe
CHESTPHENOT is trained in three stages:
- Hybrid silver supervision — combining CheXbert labels with annotations from a 72B teacher model
- Supervised fine-tuning (SFT) — on the resulting silver corpus
- Lightweight GRPO refinement — Group Relative Policy Optimization for targeted alignment
- In-distribution data
- Cross-taxonomy data
- Cross-institution data
- Over 99% of final evidence spans are locatable in the source report.
- The 3B model achieves 47.5 auditable-F1, outperforming Qwen2.5-7B one-shot prompting by 7.6 points and approaching Qwen2.5-72B.
The result is a small, locally runnable model that avoids any dependency on external inference APIs.
Evaluation
We evaluate across three human-annotated gold sets spanning:
Detection and Status
The 3B model remains below its CheXbert silver teacher in distribution, but becomes competitive under distribution shift, significantly surpassing CheXbert on cross-institution detection (+2.0 F1). Task-specific training also enables the 3B model to match or exceed substantially larger prompted models on most detection and status comparisons.
Evidence-Grounded Extraction
Key Takeaway
These results demonstrate that locally deployable models can deliver competitive, directly auditable radiology-report extraction — without relying on external inference APIs and without compromising clinical text governance.
Code and the full extraction/judge prompts will be made available at https://github.com/yukkai/ChestPheNoT.
Context
As of 2026, healthcare institutions face mounting pressure to keep protected health information within their own infrastructure while still benefiting from advances in language modeling. ChestPheNoT's evidence-span formulation directly addresses the growing demand for auditable clinical AI — a requirement now embedded in emerging regulatory frameworks for medical AI deployment — by ensuring every extracted phenotype can be traced to its exact source sentence inside the institution's firewall.
via ArXiv CL+LG
