Overview
Multimodal large language models (MLLMs) have made substantial progress in general visual question answering and cross-modal understanding. However, a pronounced evaluation gap remains for complex reasoning within the mechanical engineering domain.
Existing benchmarks predominantly target rudimentary tasks such as drawing recognition, CAD interpretation, or single-chart querying. They fall short of assessing whether models can integrate multiple images, textual conditions, physical principles, and engineering constraints to perform multi-step reasoning on authentic, intricate mechanical problems.
Introducing MechReason
To address this gap, researchers present MechReason, a benchmark derived from real mechanical engineering papers. As of 2026, it represents one of the most comprehensive domain-specific multimodal reasoning benchmarks available for engineering applications.
Scale and Composition
- 12k question-answer pairs with explicit reasoning-chain annotations
- 21k visual materials spanning nine evidence types:
- Statistical charts
- Parameter tables
- Engineering drawings
- Microscopic images
- Simulation images
- System architectures
- Real mechanical scene photos
- CAD model images
- Manufacturing flowcharts
Task Coverage
MechReason covers eight task types across four reasoning dimensions:
- Explanation
- Prediction
- Design
- Diagnosis
- Evidence extraction โ Core engineering claims are extracted, and their supporting evidence is decomposed into premises, reasoning processes, conclusions, and corroborative evidence.
- Question generation โ Shortcut-preventing questions are generated by masking posterior verification information.
- Multimodal quality validation โ Ensures task quality and the multi-hop nature of reasoning required.
- Title: MechReason: Benchmarking Multi-Image Multi-Hop Reasoning in Mechanical Engineering
- Authors: Tengyue Wang, Kang An, Chenxu Du, Zhongyu Yang, Yuanchi Zhu, Xinqi Yang, Hebao Zhu, Ziliang Wang, FaQiang Qian, Yunli Yang, Qibing Ren
- Subjects: Computer Vision and Pattern Recognition (cs.CV)
- arXiv: 2609.16012 [cs.CV]
- Submitted: 21 Aug 2026
Construction Pipeline
The benchmark is built through a rigorous four-stage pipeline:
Experimental Results
Extensive experiments demonstrate that MechReason is highly challenging. Even the most advanced models achieve only 62.89% accuracy, underscoring the difficulty of multi-image, multi-hop reasoning in this domain.
Paper Details
via ArXiv CV
