STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets
Author: Animesh Varma
arXiv: arXiv:2610.00003 [cs.CV]
Submitted: 21 May 2026
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Robotics (cs.RO)
Comments: 17 pages, 7 figures, 3 tables. Preprint
DOI: 10.48550/arXiv.2610.00003
Abstract
Vision models pretrained for frame-level appearance often struggle to infer hidden physical properties from motion. We study center-of-mass (CoM) localization for opaque, asymmetric rigid bodies from short monocular videos, where surface cues and point tracking are unreliable under self-occlusion.
We propose STATERA, which adapts a pretrained video backbone (V-JEPA) with mostly frozen weights and a lightweight temporal tubelet mixer to predict per-frame CoM heatmaps and trajectories. To support this task, we introduce the HiddenMass Benchmark, comprising 50K MuJoCo trajectories and a 63-sequence real-world test set with physically calibrated CoM ground truth.
In simulation, STATERA-50K-Sigma improves normalized CoM error from 41.7% (DINOv2) to 25.2%. In zero-shot sim-to-real transfer, we observe a fundamental trade-off in supervision: phase-aware targets can induce bimodal predictions, while phase-agnostic targets can collapse toward statistically safe centroids. Nevertheless, our phase-aware STATERA-50K-Crescent is the only evaluated method that demonstrates consistent movement toward the true hidden offset. While this leads to a monocular vector overshoot artifact that marginally increases absolute Euclidean error compared to a static geometric centroid, it improves physics capture from 2.6% to 41.0%.
These results suggest that frozen temporal representations can better separate inertial dynamics from visual geometry for hidden-parameter estimation.
Key Contributions
- STATERA framework: Adapts a pretrained V-JEPA video backbone with mostly frozen weights and a lightweight temporal tubelet mixer for per-frame CoM heatmap and trajectory prediction.
- HiddenMass Benchmark: A new benchmark comprising 50K MuJoCo trajectories and a 63-sequence real-world test set with physically calibrated CoM ground truth.
- Simulation gains: STATERA-50K-Sigma reduces normalized CoM error from 41.7% (DINOv2 baseline) to 25.2%.
- Sim-to-real insights: Identifies a supervision trade-offโphase-aware targets risk bimodal predictions, while phase-agnostic targets may collapse toward statistically safe centroids.
- Physics capture improvement: Phase-aware STATERA-50K-Crescent improves physics capture from 2.6% to 41.0%, despite a monocular vector overshoot artifact that marginally increases absolute Euclidean error.
Significance
This work suggests that frozen temporal representations can better disentangle inertial dynamics from visual geometry, offering a promising direction for hidden-parameter estimation in vision-based robotics and physical reasoning. As of 2026, as embodied AI and world models increasingly demand physically grounded perception, approaches like STATERA that bridge simulation-trained representations with zero-shot real-world deployment are becoming central to robust robotic manipulation and autonomous systems.
Citation
@article{varma2026statera,
title={STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets},
author={Varma, Animesh},
journal={arXiv preprint arXiv:2610.00003},
year={2026}
}
via ArXiv CV
