STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real

STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets


Author: Animesh Varma

arXiv: arXiv:2610.00003 [cs.CV]

Submitted: 21 May 2026

Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)

Comments: 17 pages, 7 figures, 3 tables. Preprint

DOI: 10.48550/arXiv.2610.00003


Abstract


Vision models pretrained for frame-level appearance often struggle to infer hidden physical properties from motion. We study center-of-mass (CoM) localization for opaque, asymmetric rigid bodies from short monocular videos, where surface cues and point tracking are unreliable under self-occlusion.


We propose STATERA, which adapts a pretrained video backbone (V-JEPA) with mostly frozen weights and a lightweight temporal tubelet mixer to predict per-frame CoM heatmaps and trajectories. To support this task, we introduce the HiddenMass Benchmark, comprising 50K MuJoCo trajectories and a 63-sequence real-world test set with physically calibrated CoM ground truth.


In simulation, STATERA-50K-Sigma improves normalized CoM error from 41.7% (DINOv2) to 25.2%. In zero-shot sim-to-real transfer, we observe a fundamental trade-off in supervision: phase-aware targets can induce bimodal predictions, while phase-agnostic targets can collapse toward statistically safe centroids. Nevertheless, our phase-aware STATERA-50K-Crescent is the only evaluated method that demonstrates consistent movement toward the true hidden offset. While this leads to a monocular vector overshoot artifact that marginally increases absolute Euclidean error compared to a static geometric centroid, it improves physics capture from 2.6% to 41.0%.


These results suggest that frozen temporal representations can better separate inertial dynamics from visual geometry for hidden-parameter estimation.


Key Contributions


  • STATERA framework: Adapts a pretrained V-JEPA video backbone with mostly frozen weights and a lightweight temporal tubelet mixer for per-frame CoM heatmap and trajectory prediction.
  • HiddenMass Benchmark: A new benchmark comprising 50K MuJoCo trajectories and a 63-sequence real-world test set with physically calibrated CoM ground truth.
  • Simulation gains: STATERA-50K-Sigma reduces normalized CoM error from 41.7% (DINOv2 baseline) to 25.2%.
  • Sim-to-real insights: Identifies a supervision trade-offโ€”phase-aware targets risk bimodal predictions, while phase-agnostic targets may collapse toward statistically safe centroids.
  • Physics capture improvement: Phase-aware STATERA-50K-Crescent improves physics capture from 2.6% to 41.0%, despite a monocular vector overshoot artifact that marginally increases absolute Euclidean error.

Significance


This work suggests that frozen temporal representations can better disentangle inertial dynamics from visual geometry, offering a promising direction for hidden-parameter estimation in vision-based robotics and physical reasoning. As of 2026, as embodied AI and world models increasingly demand physically grounded perception, approaches like STATERA that bridge simulation-trained representations with zero-shot real-world deployment are becoming central to robust robotic manipulation and autonomous systems.


Citation


@article{varma2026statera,
  title={STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets},
  author={Varma, Animesh},
  journal={arXiv preprint arXiv:2610.00003},
  year={2026}
}

via ArXiv CV

Related