ClinLens: A New Benchmark for Long-Horizon Clinical Data Science Agents

ai in healthcareclinical data scienceclinlenslong-horizon coding agentsmimic benchmarkmultimodal ehrprogram-first reverse synthesis

Overview


Clinical data science agents must transform heterogeneous longitudinal records into auditable analyses. However, existing benchmarks largely isolate medical question answering, structured-table reasoning, or generic scientific repositories—failing to capture the full complexity of real-world clinical workflows.


Introducing ClinLens


We present CLINLENS, a comprehensive benchmark of 200 executable tasks spanning five linked MIMIC resources:

  • Structured electronic health records
  • Clinical notes
  • Electrocardiograms
  • Chest radiographs
  • Echocardiograms

The benchmark employs a 4 × 5 taxonomy that crosses four patient-time scopes with five distinct analysis capabilities.


Methodology


A novel program-first reverse synthesis approach pairs each bounded semi-raw package with an evaluator-private reference workflow. This ensures rigorous checking of:

  • Required artifacts
  • Cohort and temporal semantics
  • Final answer correctness

Key Findings


On a fixed 126-task suite, the strongest of 24 standardized model-scaffold configurations achieved 56.3% scope-macro STRICTPASS, despite 100% EXECSUCCESS. For reference:

  • A separately configured coding agent solved 83 of 126 tasks
  • Five biomedical systems adapted to GPT-4o-mini reached at most 2.9% scope-macro STRICTPASS

These results expose a substantial gap between runnable submissions and truly correct clinical analyses, highlighting the need for more robust long-horizon reasoning in AI-driven healthcare.




Subjects: Artificial Intelligence (cs.AI)


Cite as: arXiv:2607.26155 [cs.AI]


Submission history: Submitted on 28 Jul 2026 by Jindong Han and colleagues.

via ArXiv AI

Related