ChestPheNoT: Deployable, Auditable Label-Status-Evidence

ChestPheNoT: Deployable, Auditable Label-Status-Evidence Extraction from Radiology Reports


Authors: Kai Yu, Chenyu Zhu, Zaifu Zhan, Meijia Song, Min Zeng, Xiaoyi Chen, Mingquan Lin, Rui Zhang


Venue: Accepted at IEEE Healthcom 2026


Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)


arXiv: arXiv:2609.31629 [cs.CL]


Code: https://github.com/yukkai/ChestPheNoT




Overview


Structured phenotype extraction from radiology reports underpins cohort construction, quality auditing, and clinical analytics. However, practical deployment imposes two hard constraints: inference must run locally, and predictions must be auditable. Expert annotations, meanwhile, remain scarce.


Existing approaches fall short of these requirements. Conventional labelers output structured findings and assertion states but provide no supporting evidence. API-hosted large language models, by contrast, may be unusable when clinical text is prohibited from leaving institutional infrastructure.


Introducing ChestPheNoT


We present CHESTPHENOT, a compact 0.5–3B parameter language model that jointly extracts:


  • Finding labels — the phenotype of interest
  • Three-class status — present, absent, or uncertain
  • Verbatim supporting evidence spans — the exact source text justifying each prediction

This joint formulation makes every prediction directly traceable to the report text, enabling auditing without a separate explanation step.


Training Recipe


CHESTPHENOT is trained in three stages:


  1. Hybrid silver supervision — combining CheXbert labels with annotations from a 72B teacher model
  2. Supervised fine-tuning (SFT) — on the resulting silver corpus
  3. Lightweight GRPO refinement — Group Relative Policy Optimization for targeted alignment

  4. The result is a small, locally runnable model that avoids any dependency on external inference APIs.


    Evaluation


    We evaluate across three human-annotated gold sets spanning:


    • In-distribution data
    • Cross-taxonomy data
    • Cross-institution data

    Detection and Status


    The 3B model remains below its CheXbert silver teacher in distribution, but becomes competitive under distribution shift, significantly surpassing CheXbert on cross-institution detection (+2.0 F1). Task-specific training also enables the 3B model to match or exceed substantially larger prompted models on most detection and status comparisons.


    Evidence-Grounded Extraction


    • Over 99% of final evidence spans are locatable in the source report.
    • The 3B model achieves 47.5 auditable-F1, outperforming Qwen2.5-7B one-shot prompting by 7.6 points and approaching Qwen2.5-72B.

    Key Takeaway


    These results demonstrate that locally deployable models can deliver competitive, directly auditable radiology-report extraction — without relying on external inference APIs and without compromising clinical text governance.


    Code and the full extraction/judge prompts will be made available at https://github.com/yukkai/ChestPheNoT.


    Context


    As of 2026, healthcare institutions face mounting pressure to keep protected health information within their own infrastructure while still benefiting from advances in language modeling. ChestPheNoT's evidence-span formulation directly addresses the growing demand for auditable clinical AI — a requirement now embedded in emerging regulatory frameworks for medical AI deployment — by ensuring every extracted phenotype can be traced to its exact source sentence inside the institution's firewall.

    via ArXiv CL+LG

Related