UltraBench 2: A New Benchmark for Evaluating Vision Foundation

UltraBench 2: Towards Robust Evaluation of Vision Foundation Models on Ultrasound


Authors: Ashwath Radhachandran, Adam Tupper, Christian Gagné, William Speier


Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)


Submitted: 23 September 2026


arXiv: 2609.28610 [cs.CV]


Abstract


Benchmarking has become an increasingly critical component of machine learning research and the domains where it is applied, including healthcare. Yet, despite the steady development of new ultrasound foundation models in recent years, the creation of well-designed benchmarks to evaluate them has lagged behind. This deficiency has led to fragmented and inconsistent evaluations of competing models, making it difficult to measure genuine progress in the field.


To address this issue, we introduce UltraBench 2, a comprehensive benchmark with wide anatomical and task coverage, and a focus on standardization, reproducibility, and ease-of-use. Using this benchmark, we compare existing vision foundation models for ultrasound image analysis. Our analyses demonstrate that ultrasound-specific pretraining still leads on classification tasks, but that state-of-the-art general-purpose models have drawn level on segmentation.


Key Contributions


  • Comprehensive coverage: UltraBench 2 spans a wide range of anatomical regions and clinical tasks, addressing the fragmentation that has hindered meaningful comparison across studies.
  • Standardization and reproducibility: The benchmark emphasizes consistent evaluation protocols and reproducible results, allowing researchers to track progress reliably.
  • Comparative evaluation: The study benchmarks existing vision foundation models for ultrasound, revealing that domain-specific pretraining maintains an edge in classification, while top general-purpose models now match performance on segmentation.

Why It Matters


As ultrasound foundation models proliferate in 2026, the absence of rigorous, standardized benchmarks has made it difficult to compare competing approaches. UltraBench 2 aims to fill this gap, providing the medical imaging and computer vision communities with a robust evaluation framework to measure and accelerate progress.


Access


via ArXiv CV

Related