MechReason: A New Benchmark for Multi-Image, Multi-Hop Reasoning

Overview


Multimodal large language models (MLLMs) have made substantial progress in general visual question answering and cross-modal understanding. However, a pronounced evaluation gap remains for complex reasoning within the mechanical engineering domain.


Existing benchmarks predominantly target rudimentary tasks such as drawing recognition, CAD interpretation, or single-chart querying. They fall short of assessing whether models can integrate multiple images, textual conditions, physical principles, and engineering constraints to perform multi-step reasoning on authentic, intricate mechanical problems.


Introducing MechReason


To address this gap, researchers present MechReason, a benchmark derived from real mechanical engineering papers. As of 2026, it represents one of the most comprehensive domain-specific multimodal reasoning benchmarks available for engineering applications.


Scale and Composition


  • 12k question-answer pairs with explicit reasoning-chain annotations
  • 21k visual materials spanning nine evidence types:
  • Statistical charts
  • Parameter tables
  • Engineering drawings
  • Microscopic images
  • Simulation images
  • System architectures
  • Real mechanical scene photos
  • CAD model images
  • Manufacturing flowcharts

Task Coverage


MechReason covers eight task types across four reasoning dimensions:


  1. Explanation
  2. Prediction
  3. Design
  4. Diagnosis

  5. Construction Pipeline


    The benchmark is built through a rigorous four-stage pipeline:


    1. Evidence extraction โ€” Core engineering claims are extracted, and their supporting evidence is decomposed into premises, reasoning processes, conclusions, and corroborative evidence.
    2. Question generation โ€” Shortcut-preventing questions are generated by masking posterior verification information.
    3. Multimodal quality validation โ€” Ensures task quality and the multi-hop nature of reasoning required.

    4. Experimental Results


      Extensive experiments demonstrate that MechReason is highly challenging. Even the most advanced models achieve only 62.89% accuracy, underscoring the difficulty of multi-image, multi-hop reasoning in this domain.


      Paper Details


      • Title: MechReason: Benchmarking Multi-Image Multi-Hop Reasoning in Mechanical Engineering
      • Authors: Tengyue Wang, Kang An, Chenxu Du, Zhongyu Yang, Yuanchi Zhu, Xinqi Yang, Hebao Zhu, Ziliang Wang, FaQiang Qian, Yunli Yang, Qibing Ren
      • Subjects: Computer Vision and Pattern Recognition (cs.CV)
      • arXiv: 2609.16012 [cs.CV]
      • Submitted: 21 Aug 2026

      via ArXiv CV

Related