Emergent Object Binding Has a Finite Spatial Horizon

Emergent Object Binding Has a Finite Spatial Horizon


Authors: Mayank Singal

arXiv: 2610.00006 [cs.CV]

Submitted: 9 July 2026

Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

Comments: 14 pages, 4 figures

DOI: 10.48550/arXiv.2610.00006


Abstract


Pretrained Vision Transformers encode whether two image patches belong to the same object. This IsSameObject signal is decodable from frozen patch embeddings at high accuracy, which suggests that object binding emerges from self-supervised pretraining alone. We show that this single accuracy number hides the structure of the signal.


Binding is local: the probability that two patches of the same object are decoded as bound falls off monotonically with the distance between them and levels off at a nonzero floor β€” a falloff well described by an exponential with a finite length scale. This decay holds across object sizes, across three families of probe, on both ADE20K and COCO, and across DINO and CLIP backbones, indicating that it is a property of the representation rather than of the decoder.


Reading binding as local spatial coherence with a finite range accounts for a set of behaviors that the aggregate score leaves unexplained:


  • Binding weakens on large objects.
  • Distinct objects of the same class are separated less reliably than objects of different classes.
  • Object parts are grouped with their wholes.

By contrast, binding is unaffected by occlusion once object size is controlled. We map each behavior with confounds controlled.


As a preliminary observation, the horizon and its floor are organized at different depths in DINOv2 and DINOv3, which we report as suggestive given the small number of layers probed and the confound between the two models.


Key Takeaways


  • Emergent binding is local, not global. Frozen Vision Transformer embeddings encode same-object relationships, but the probability of correct binding decays exponentially with patch distance β€” levelling off at a nonzero floor rather than dropping to zero.
  • Robust across setups. The decay pattern is consistent across object sizes, three probe families, ADE20K and COCO datasets, and both DINO and CLIP backbones, confirming it as a property of the learned representation.
  • Explains known failure modes. The finite-horizon view accounts for weakened binding on large objects, unreliable separation of same-class instances, and part–whole grouping β€” while showing binding is occlusion-invariant once object size is controlled.
  • Layer-depth differences. Preliminary findings suggest the spatial horizon and its floor are organized at different depths in DINOv2 versus DINOv3, though the small number of layers probed and cross-model confounds warrant caution.

Submission History


  • [v1] Thu, 9 Jul 2026 02:41:15 UTC (1,052 KB)

via ArXiv CV

Related