Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning
1ACI PLC, Bangladesh 2University of Dhaka 3BRAC University 4QCRI, Qatar 5Amazon GenAI, USA
arXiv preprint arXiv:2609.33200, 2026
* Equal supervision · ‡ Work done outside role at Amazon
TL;DR
A privileged teacher that has seen the verified solution knows not only what token comes next but also which earlier steps matter. OPASD distills that attention too, after projecting it onto the context the student can actually see, and gets more accurate, shorter, and cheaper reasoning.
- +8.40pt Avg@12 over OPSD on Qwen3-4B (+4.98 to +8.40 across 3 sizes)
- −73.9%generated rollout tokens vs. token-only OPSD
- −72.6%estimated training compute (EFLOPs)
- 1.53×faster wall-clock training

Abstract
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53× faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.
The missing signal
Motivation
In on-policy self-distillation (OPSD), the same model plays two roles. As a student it sees only the problem and its own partial answer; as a privileged teacher it also sees a verified solution. The teacher's next-token distribution is dense, per-token feedback on the student's own trajectory.
But next-token distributions say what to predict, not where to look. In multi-step reasoning, later steps hinge on constraints and intermediate results established much earlier, and the teacher's attention over that context is a solution-conditioned signal that OPSD leaves unused. Distilling it is not trivial: teacher and student attend over different contexts, because only the teacher can attend to solution tokens.
Method
Four steps per training iteration

- On-policy samplingThe student generates its own reasoning trajectory for a problem x, so supervision is always about states the student actually visits.
- Two views, one trajectoryThe teacher re-reads the same prefix with the verified solution y* in context. Both models emit next-token and final-layer attention distributions.
- Student-support projectionAttention mass on solution-only positions is removed and the remainder renormalized, giving a target defined entirely over context the student can see.
- Joint distillationThe student minimizes L = Ltok + λattn Lattn, using JSD for attention, so it learns what to predict and where to look.
Results
Qwen3 1.7B / 4B / 8B · Avg@12 accuracy (%)
OPASD is the best method at every model size on all four competition-level benchmarks. Token-only OPSD is unstable at this scale: on Qwen3-4B it falls below the base model, while OPASD improves on it by 8.40 points.
| Method | AIME24 | AIME25 | AIME26 | HMMT25 | Average |
|---|---|---|---|---|---|
| Qwen3-8B | |||||
| Base | 60.83 | 48.89 | 51.39 | 29.17 | 47.57 |
| OPSD | 59.16 | 51.97 | 53.33 | 33.33 | 49.45 |
| OPASD | 67.50 | 57.22 | 61.11 | 35.83 | 55.42 |
| Qwen3-4B | |||||
| Base | 58.33 | 47.50 | 51.11 | 30.27 | 46.80 |
| OPSD | 51.11 | 43.88 | 46.11 | 28.33 | 42.36 |
| OPASD | 63.33 | 50.27 | 55.56 | 33.89 | 50.76 |
| Qwen3-1.7B | |||||
| Base | 33.61 | 30.28 | 31.94 | 19.17 | 28.75 |
| OPSD | 36.70 | 28.33 | 32.78 | 18.33 | 29.04 |
| OPASD | 43.05 | 35.27 | 34.44 | 23.33 | 34.02 |
More accurate and cheaper
The extra attention objective does not add cost; it removes it. Because OPASD avoids response-length inflation, rollouts are far shorter, so the whole run is faster and lighter.

Stable training dynamics

What matters
Ablations on Qwen3-1.7B · average Avg@12
- Divergence: JSD (32.85) > reverse KL (32.36) > forward KL (31.73). A symmetric, bounded objective balances mode-covering and mode-seeking behaviour.
- Attention weight: λattn = 0.5 is best (34.02), ahead of 1.0 (32.85) and 2.0 (31.82).
- Depth: distilling only the final layer (34.02) beats the final two (32.54) or three (31.04) layers.
Citation
@article{arib2026opasd,
title = {Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning},
author = {Arib, Safaeid Hossain and Akter, Rabeya and Swapnil, Ismam Nur and
Sayeedi, Md. Faiyaz Abdullah and Mohiuddin, Tasnim and Islam, Md Mofijul},
journal = {arXiv preprint arXiv:2609.33200},
year = {2026}
}