Overview
Clinical data science agents must transform heterogeneous longitudinal records into auditable analyses. However, existing benchmarks largely isolate medical question answering, structured-table reasoning, or generic scientific repositories—failing to capture the full complexity of real-world clinical workflows.
Introducing ClinLens
We present CLINLENS, a comprehensive benchmark of 200 executable tasks spanning five linked MIMIC resources:
- Structured electronic health records
- Clinical notes
- Electrocardiograms
- Chest radiographs
- Echocardiograms
The benchmark employs a 4 × 5 taxonomy that crosses four patient-time scopes with five distinct analysis capabilities.
Methodology
A novel program-first reverse synthesis approach pairs each bounded semi-raw package with an evaluator-private reference workflow. This ensures rigorous checking of:
- Required artifacts
- Cohort and temporal semantics
- Final answer correctness
Key Findings
On a fixed 126-task suite, the strongest of 24 standardized model-scaffold configurations achieved 56.3% scope-macro STRICTPASS, despite 100% EXECSUCCESS. For reference:
- A separately configured coding agent solved 83 of 126 tasks
- Five biomedical systems adapted to GPT-4o-mini reached at most 2.9% scope-macro STRICTPASS
These results expose a substantial gap between runnable submissions and truly correct clinical analyses, highlighting the need for more robust long-horizon reasoning in AI-driven healthcare.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2607.26155 [cs.AI]
Submission history: Submitted on 28 Jul 2026 by Jindong Han and colleagues.
via ArXiv AI
