PalmLeaf-VQA: A Multi-Script Visual Question Answering Benchmark

PalmLeaf-VQA: A Multi-Script Visual Question Answering Benchmark for Historical Palm-Leaf Manuscripts


Authors: Nimol Thuon, Jun Du, Panhapin Theang

Submitted: 14 Sep 2026

arXiv: 2609.31651 [cs.CV]

DOI: 10.48550/arXiv.2609.31651


Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)


Overview


Historical manuscripts remain largely absent from modern vision-language benchmarks, leaving open the question of how well multimodal large language models (MLLMs) handle culturally diverse, degraded, and non-Latin document images. To address this gap, the authors introduce PalmLeaf-VQA, a multi-script visual question answering benchmark for historical palm-leaf manuscript understanding across South and Southeast Asian traditions.


Benchmark Composition


PalmLeaf-VQA comprises:


  • 923 curated manuscript images
  • 7,384 question–answer pairs

These are drawn from eight collection groups:


  • Balinese
  • Grantha
  • Jathakam
  • Kambaramayanam
  • Kannada
  • Khmer
  • Sundanese
  • Tamil

Unlike recognition-oriented resources, the benchmark targets manuscript-aware visual reasoning over preservation-relevant cues, including physical condition, line structure, material and coating, binding holes, margins, symbols, drawings, and localized visual artifacts.


Evaluation and Findings


The authors evaluate recent proprietary and open-weight MLLMs under both open-answer and constrained-answer prompting, and provide fine-grained analysis across collections, question categories, and task types. The strongest evaluated model reaches only 58.00% exact-match accuracy on the held-out test split, revealing substantial limitations in current MLLMs for rare-script, degraded-layout, and preservation-oriented document understanding.


Significance


As of 2026, MLLMs have achieved strong performance on mainstream document and chart understanding tasks, yet their capabilities remain uneven across low-resource scripts and culturally specific artifacts. PalmLeaf-VQA provides a standardized benchmark for advancing culturally grounded and layout-aware multimodal document analysis, and it highlights a persistent performance gap that future work in document AI and cultural heritage preservation must close.

via ArXiv CV

Related