PalmLeaf-VQA: A Multi-Script Visual Question Answering Benchmark for Historical Palm-Leaf Manuscripts
Authors: Nimol Thuon, Jun Du, Panhapin Theang
Submitted: 14 Sep 2026
arXiv: 2609.31651 [cs.CV]
DOI: 10.48550/arXiv.2609.31651
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Overview
Historical manuscripts remain largely absent from modern vision-language benchmarks, leaving open the question of how well multimodal large language models (MLLMs) handle culturally diverse, degraded, and non-Latin document images. To address this gap, the authors introduce PalmLeaf-VQA, a multi-script visual question answering benchmark for historical palm-leaf manuscript understanding across South and Southeast Asian traditions.
Benchmark Composition
PalmLeaf-VQA comprises:
- 923 curated manuscript images
- 7,384 questionβanswer pairs
These are drawn from eight collection groups:
- Balinese
- Grantha
- Jathakam
- Kambaramayanam
- Kannada
- Khmer
- Sundanese
- Tamil
Unlike recognition-oriented resources, the benchmark targets manuscript-aware visual reasoning over preservation-relevant cues, including physical condition, line structure, material and coating, binding holes, margins, symbols, drawings, and localized visual artifacts.
Evaluation and Findings
The authors evaluate recent proprietary and open-weight MLLMs under both open-answer and constrained-answer prompting, and provide fine-grained analysis across collections, question categories, and task types. The strongest evaluated model reaches only 58.00% exact-match accuracy on the held-out test split, revealing substantial limitations in current MLLMs for rare-script, degraded-layout, and preservation-oriented document understanding.
Significance
As of 2026, MLLMs have achieved strong performance on mainstream document and chart understanding tasks, yet their capabilities remain uneven across low-resource scripts and culturally specific artifacts. PalmLeaf-VQA provides a standardized benchmark for advancing culturally grounded and layout-aware multimodal document analysis, and it highlights a persistent performance gap that future work in document AI and cultural heritage preservation must close.
via ArXiv CV
