Improving OCR Faithfulness via Gated and Attenuated On-Policy

Overview


Vision-language models (VLMs) may rewrite anomalous text in images into linguistically plausible expressions, compromising the faithfulness of OCR transcription. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improvesβ€”both across training checkpoints and across response groups with different task rewards.


Method: GAD-RL


Motivated by this observation, the authors introduce GAD-RL, which adaptively regulates teacher supervision during joint post-training according to the student's current task performance and local distributions. A frozen teacher conditions on reference transcriptions and student-generated prefixes. GAD-RL comprises three key mechanisms:


  • Gating: Disables distillation for response groups containing an output with task reward of at least 0.95.
  • Attenuation: Continuously attenuates distillation strength as group-mean reward increases.
  • Local weighting: Weights forward KL divergence by the student's probability of the teacher's Top-1 token, moderating local auxiliary updates when student support for that candidate is low.

Results


On Qwen3.5-2B, GAD-RL achieves:


  • 59.92% Micro Recall on CHAOS-Bench, surpassing GRPO and GRPO+OPD (fixed-weight) by 8.45 and 4.43 percentage points, respectively.
  • 91.18 Overall score on OmniDocBench v1.6.

Paper Details


  • arXiv: 2609.38282 [cs.AI]
  • Submitted: 29 September 2026
  • Authors: Baode Wang, Zuming Huang, Kexuan Ren, Jun Huang, Wei Chu
  • Subject: Artificial Intelligence (cs.AI)
  • DOI: 10.48550/arXiv.2609.38282

via ArXiv AI

Related