Computer Vision and Pattern Recognition (cs.CV) — arXiv:2607.21722
Published: 23 July 2026
Authors: Liqiang Jing, Xiong Zhou, Siddharth Varia, Neha Anna John, Xinya Du, and Vassilis N. Ioannidis
Abstract
Large Vision-Language Models (LVLMs) demonstrate strong perceptual capabilities, yet they remain vulnerable in complex visual reasoning tasks. Existing benchmarks predominantly focus on symbolic mathematical or scientific problems and simplistic vision-centric tasks, offering limited assessment of complex visual reasoning and, critically, logical consistency—a fundamental requirement for reliable reasoning systems.
To address this gap, we introduce ConVBench, a vision-centric reasoning benchmark designed to evaluate logical consistency. In this benchmark, each image is paired with two logically equivalent questions spanning six categories: action and state, complex counting, spatial reasoning, causal and intent understanding, commonsense reasoning, and temporal perception. We further propose two complementary evaluation metrics—logical consistency and robust accuracy—that jointly assess both the correctness and the consistency of model responses.
In addition, we present ConVLM, a reinforcement learning framework that enhances LVLM reasoning through Group Relative Policy Optimization (GRPO) with a novel consistency reward. ConVLM leverages automatically generated logically equivalent question-answer pairs and a dual-reward design that combines accuracy-based and consistency-based signals. This approach encourages agreement between paired responses and operates effectively with or without strict answer supervision.
2026 Update: With the rapid advancement of LVLMs toward deployment in high-stakes applications such as autonomous driving and medical imaging, ensuring robust and logically consistent reasoning has become critical. ConVBench and ConVLM provide a comprehensive toolkit for evaluating and improving consistency in vision-language models, addressing a key bottleneck in current AI systems.
via ArXiv CV
