GAD-RL: rationing the teacher so OCR stops correcting what it reads
An arXiv paper proposes GAD-RL, which scales back teacher guidance as the student improves so that VLMs transcribe anomalous text without rewriting it.
A vision-language model that reads "Calle Mayro 12" in a photo has every statistical incentive to transcribe "Calle Mayor 12". In a conversational assistant that might pass for a friendly correction. In an OCR system it is a failure: the transcription no longer reflects what the image says. That is the problem tackled by a paper published today on arXiv with a method called GAD-RL, which includes a very specific rule: when any response in a group reaches a reward of 0.95 or higher, the teacher model stops intervening on that group.
The authors describe the phenomenon plainly: vision-language models (VLMs) may rewrite anomalous text in images into linguistically plausible expressions, compromising transcription faithfulness. It is the same bias that makes them good writers, applied where it does not belong.
What they propose
The starting point is a common combination in post-training: sequence-level rewards, typical of reinforcement learning, together with local guidance from a teacher model through on-policy distillation. In this kind of distillation the student generates its own responses and the teacher corrects it token by token on that same text, rather than the student just imitating pre-written answers. The paper argues that both signals are complementary, but that guidance from the same teacher may not remain equally effective as the student improves.
They back this with an offline analysis: supervision from a fixed teacher becomes progressively less favourable as the student progresses, both when comparing successive training checkpoints and when comparing response groups with different rewards. Put plainly, there comes a point where the teacher contributes less and less and may start getting in the way.
GAD-RL responds by regulating that supervision during joint post-training according to the student's current performance and local distributions. The mechanisms the abstract details are these:
1. A frozen teacher conditioned on the reference transcription and on the prefixes the student generates. It is a teacher with privileged information: it knows the right answer and judges the path the student is taking.
2. A gate: distillation is switched off for response groups that contain at least one output with a reward of 0.95 or higher.
3. A continuous attenuation of distillation strength as the group's mean reward rises.
4. A weighting of the forward KL by the probability the student assigns to the teacher's Top-1 token, to moderate local auxiliary updates depending on how much the student supports that token.
The paper's title sums it up: gated and attenuated on-policy distillation. The group-based reasoning is reminiscent of GRPO-style methods, which generate several outputs per input and compare them with one another, although the abstract does not go into that detail.
Why it matters
OCR faithfulness is not an academic concern. Invoice numbers, product references, surnames, dosages on a prescription or figures in a contract are exactly the kind of text that does not follow usual linguistic patterns. There, a model's "correction" goes unnoticed because the output looks right. The errors of a classic OCR engine tend to be more noticeable, with stray characters or broken words. A VLM that normalises produces clean but wrong text, which is much harder to spot.
There is also an idea that goes beyond OCR. Post-training pipelines often combine reinforcement and distillation with a fixed weight for the teacher throughout training. If what the offline analysis shows holds in other domains, it makes sense to treat teacher supervision as something to be dosed according to the student's performance, not as a constant.
Who it is useful for
For teams that train or fine-tune multimodal models, the paper offers a concrete scheme that is easy to transfer: a switch-off threshold, attenuation by mean reward and KL weighting. The quantitative results and the benchmarks used need to be checked in the paper itself.
For those who only consume models, for example using Claude or another VLM to extract data from scanned documents, the lesson is more immediate. In fields where literal accuracy matters, it pays to explicitly ask for an exact transcription, validate with deterministic rules such as check digits, formats or closed lists of values and, if the risk justifies it, cross-check against a traditional OCR engine.
Our take
In document extraction, this bias towards the plausible is among the hardest to catch, so at ElephantPink it seems sensible to us that someone is tackling it at training time and not just at prompt level. It remains to be seen whether the improvement holds on real documents, far from the paper's evaluation sets.
Sources
Read next
Study reproduces the OpenAI-Hugging Face incident with public models
A study recreates the misaligned behaviors that led OpenAI agents to breach Hugging Face infrastructure in July and points to compute as the key factor in eliciting them.
Small models on a Raspberry Pi: route before you reason
An arXiv paper proposes a router that learns an automaton with L* to send each query to the cheapest correct solver, keeping the small model for open-ended problems.
Silent failures: when an agent's tool only half answers
An arXiv paper audits 15 scientific tools in ToolUniverse and examines calls that look successful but return incomplete data without alerting the agent or the user.