arXiv 2026 New LLM ReasoningPost-TrainingDistillation

Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

Safaeid Hossain Arib1, Rabeya Akter2, Ismam Nur Swapnil1, Md. Faiyaz Abdullah Sayeedi3, Tasnim Mohiuddin4,*, Md Mofijul Islam5,*,‡

1ACI PLC, Bangladesh   2University of Dhaka   3BRAC University   4QCRI, Qatar   5Amazon GenAI, USA

arXiv preprint arXiv:2609.33200, 2026

* Equal supervision  ·  ‡ Work done outside role at Amazon

TL;DR

A privileged teacher that has seen the verified solution knows not only what token comes next but also which earlier steps matter. OPASD distills that attention too, after projecting it onto the context the student can actually see, and gets more accurate, shorter, and cheaper reasoning.

  • +8.40pt Avg@12 over OPSD on Qwen3-4B (+4.98 to +8.40 across 3 sizes)
  • −73.9%generated rollout tokens vs. token-only OPSD
  • −72.6%estimated training compute (EFLOPs)
  • 1.53×faster wall-clock training
Overview figure for Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning
OPSD vs. OPASD. (a) OPSD aligns teacher and student next-token distributions: what to predict. (b) OPASD additionally aligns projected attention distributions: where to look.

Abstract

On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53× faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.

The missing signal

Motivation

In on-policy self-distillation (OPSD), the same model plays two roles. As a student it sees only the problem and its own partial answer; as a privileged teacher it also sees a verified solution. The teacher's next-token distribution is dense, per-token feedback on the student's own trajectory.

But next-token distributions say what to predict, not where to look. In multi-step reasoning, later steps hinge on constraints and intermediate results established much earlier, and the teacher's attention over that context is a solution-conditioned signal that OPSD leaves unused. Distilling it is not trivial: teacher and student attend over different contexts, because only the teacher can attend to solution tokens.

Method

Four steps per training iteration

OPASD method overview: on-policy sampling, teacher-student distributions, student-support projection, and joint distillation
OPASD overview. The student samples a rollout; teacher and student produce token and attention distributions over the same trajectory; the teacher's attention is projected onto student-visible positions and renormalized; and both signals are distilled jointly.
  1. On-policy samplingThe student generates its own reasoning trajectory for a problem x, so supervision is always about states the student actually visits.
  2. Two views, one trajectoryThe teacher re-reads the same prefix with the verified solution y* in context. Both models emit next-token and final-layer attention distributions.
  3. Student-support projectionAttention mass on solution-only positions is removed and the remainder renormalized, giving a target defined entirely over context the student can see.
  4. Joint distillationThe student minimizes L = Ltok + λattn Lattn, using JSD for attention, so it learns what to predict and where to look.

Results

Qwen3 1.7B / 4B / 8B · Avg@12 accuracy (%)

OPASD is the best method at every model size on all four competition-level benchmarks. Token-only OPSD is unstable at this scale: on Qwen3-4B it falls below the base model, while OPASD improves on it by 8.40 points.

MethodAIME24AIME25AIME26HMMT25Average
Qwen3-8B
Base60.8348.8951.3929.1747.57
OPSD59.1651.9753.3333.3349.45
OPASD67.5057.2261.1135.8355.42
Qwen3-4B
Base58.3347.5051.1130.2746.80
OPSD51.1143.8846.1128.3342.36
OPASD63.3350.2755.5633.8950.76
Qwen3-1.7B
Base33.6130.2831.9419.1728.75
OPSD36.7028.3332.7818.3329.04
OPASD43.0535.2734.4423.3334.02

More accurate and cheaper

The extra attention objective does not add cost; it removes it. Because OPASD avoids response-length inflation, rollouts are far shorter, so the whole run is faster and lighter.

Accuracy versus wall-clock: OPASD reaches higher Avg@12 in 5.27 hours versus 8.04 hours for OPSD
OPASD Pareto-dominates OPSD on Qwen3-4B. Higher Avg@12 with lower wall-clock time, estimated compute, generated tokens (bubble area), and peak GPU memory (−11.9%).

Stable training dynamics

Training and validation accuracy, response length, and epistemic token usage over training for OPSD and OPASD
Training dynamics on Qwen3-4B. OPASD keeps training and validation accuracy stable with concise responses. OPSD shows declining accuracy, response-length inflation, and much heavier use of uncertainty and reconsideration markers.

What matters

Ablations on Qwen3-1.7B · average Avg@12

  • Divergence: JSD (32.85) > reverse KL (32.36) > forward KL (31.73). A symmetric, bounded objective balances mode-covering and mode-seeking behaviour.
  • Attention weight: λattn = 0.5 is best (34.02), ahead of 1.0 (32.85) and 2.0 (31.82).
  • Depth: distilling only the final layer (34.02) beats the final two (32.54) or three (31.04) layers.

Citation

@article{arib2026opasd,
  title   = {Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning},
  author  = {Arib, Safaeid Hossain and Akter, Rabeya and Swapnil, Ismam Nur and
             Sayeedi, Md. Faiyaz Abdullah and Mohiuddin, Tasnim and Islam, Md Mofijul},
  journal = {arXiv preprint arXiv:2609.33200},
  year    = {2026}
}