Overview
Vision-language models (VLMs) may rewrite anomalous text in images into linguistically plausible expressions, compromising the faithfulness of OCR transcription. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improvesβboth across training checkpoints and across response groups with different task rewards.
Method: GAD-RL
Motivated by this observation, the authors introduce GAD-RL, which adaptively regulates teacher supervision during joint post-training according to the student's current task performance and local distributions. A frozen teacher conditions on reference transcriptions and student-generated prefixes. GAD-RL comprises three key mechanisms:
- Gating: Disables distillation for response groups containing an output with task reward of at least 0.95.
- Attenuation: Continuously attenuates distillation strength as group-mean reward increases.
- Local weighting: Weights forward KL divergence by the student's probability of the teacher's Top-1 token, moderating local auxiliary updates when student support for that candidate is low.
Results
On Qwen3.5-2B, GAD-RL achieves:
- 59.92% Micro Recall on CHAOS-Bench, surpassing GRPO and GRPO+OPD (fixed-weight) by 8.45 and 4.43 percentage points, respectively.
- 91.18 Overall score on OmniDocBench v1.6.
Paper Details
- arXiv: 2609.38282 [cs.AI]
- Submitted: 29 September 2026
- Authors: Baode Wang, Zuming Huang, Kexuan Ren, Jun Huang, Wei Chu
- Subject: Artificial Intelligence (cs.AI)
- DOI: 10.48550/arXiv.2609.38282
via ArXiv AI
