Computer Science > Computer Vision and Pattern Recognition
arXiv:2610.10563 (cs.CV) — Submitted on 1 Oct 2026
Title
SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows
Authors: Albert Gao, Bing Xue, Andrea Zanette
Comments: Accepted by NeurIPS 2026. Project page available at https://bogao-code.github.io/SLVR/
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.10563 [cs.CV] (or arXiv:2610.10563v1 [cs.CV] for this version)
Abstract
Multimodal large language models (MLLMs) often answer visual reasoning questions by relying on linguistic priors rather than task-relevant visual evidence. Textual chain-of-thought reasoning can partially mitigate this issue by encouraging models to decompose visual questions into intermediate evidence-seeking steps, but generating these steps autoregressively increases inference cost. Latent reasoning avoids explicit rationale generation, but existing approaches provide limited control over what intermediate states encode, making it difficult to impose separate supervision for planning, grounding, and evidence selection.
We propose Structured Latent Visual Reasoning (SLVR), a training framework that bridges explicit chain-of-thought and latent reasoning by organizing multimodal reasoning into typed latent stages for planning, grounding, evidence selection, and reasoning integration.
SLVR first trains the model to rely on the image by masking answer-revealing text and contrasting the correct answer with visually plausible distractors. It then organizes reasoning into latent stages for planning, grounding, evidence selection, and integration, supervising each stage with the corresponding signal: plans, boxes, visual evidence, and final rationales. This gives latent reasoning an explicit functional structure while avoiding generated textual chains at inference time.
Built on Qwen2.5-VL-7B, SLVR improves consistently across multimodal reasoning benchmarks, with absolute gains of +9.4 on MMVP and +14.2 on BLINK Relation, as well as improvements on V\*, MathVista, and ChartQA. These results suggest that structured latent supervision can improve fine-grained visual reasoning without the decoding overhead of textual CoT.
Project page: https://bogao-code.github.io/SLVR/
Key Contributions
- Structured latent reasoning framework: Organizes multimodal reasoning into typed latent stages — planning, grounding, evidence selection, and reasoning integration — providing explicit functional structure to latent states.
- Image-grounded training: Masks answer-revealing text and contrasts correct answers against visually plausible distractors, forcing the model to rely on visual evidence rather than linguistic priors.
- Stage-specific supervision: Each latent stage is supervised with a corresponding signal (plans, boxes, visual evidence, final rationales), enabling separate control over intermediate representations.
- Inference efficiency: Avoids autoregressive generation of textual chains of thought, reducing decoding overhead while preserving reasoning quality.
Experimental Results
SLVR, built on Qwen2.5-VL-7B, demonstrates consistent improvements across multiple multimodal reasoning benchmarks:
| Benchmark | Improvement |
| --- | --- |
| MMVP | +9.4 |
| BLINK Relation | +14.2 |
| V\* | Improvement |
| MathVista | Improvement |
| ChartQA | Improvement |
These results indicate that structured latent supervision can enhance fine-grained visual reasoning without incurring the decoding overhead associated with textual chain-of-thought methods.
via ArXiv CV
