Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning
1ACI PLC, Bangladesh 2University of Dhaka 3BRAC University 4QCRI, Qatar 5Amazon GenAI, USA
arXiv preprint arXiv:2609.33200, 2026
* Equal supervision · ‡ Work done outside role at Amazon
TL;DR
A privileged teacher that has seen the verified solution knows not only what token comes next but also which earlier steps matter. OPASD distills that attention too, after projecting it onto the context the student can actually see, and gets more accurate, shorter, and cheaper reasoning.
- +8.40pt Avg@12 over OPSD on Qwen3-4B (+4.98 to +8.40 across 3 sizes)
- −73.9%generated rollout tokens vs. token-only OPSD
- −72.6%estimated training compute (EFLOPs)
- 1.53×faster wall-clock training

Abstract
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53× faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.
What to predict is not where to look
Motivation
Reinforcement learning with verifiable rewards gives a reasoning model one number per response. On-policy distillation is denser: a teacher scores every prefix of the student's own trajectory with a full next-token distribution. On-policy self-distillation (OPSD) goes further and needs no separate teacher. The same model plays both roles, and the "teacher" copy is simply shown the verified solution.
That privileged teacher knows more than the next token. Multi-step reasoning hinges on constraints and intermediate results established much earlier, and the teacher's attention shows which of them matter at each step. OPSD discards this signal. OPASD distills it.
The catch is a support mismatch
Next-token distributions are directly comparable because teacher and student share one vocabulary. Attention distributions are not. They live over context positions, and the teacher's context contains solution tokens that the student never sees. Some of the teacher's attention mass sits on positions that do not exist for the student, so the two distributions cannot be aligned as they are.
Method
On-policy sampling, two views, projection, joint distillation

1. On-policy self-distillation
One language model defines two policies. The teacher is the frozen initial checkpoint conditioned on the problem and its verified solution, and the student conditions on the problem alone, exactly as at inference time.
The student samples its own trajectory \(\hat{y} \sim p_S(\cdot \mid x)\), and the teacher scores that same trajectory token by token rather than generating its own. The token-level objective matches their next-token distributions at every step.
2. Attention from both views
For each generated token, OPASD reads the head-averaged final-layer attention of both models. The student attends over its visible support \(V_t^S\) (problem and prefix). The teacher attends over \(V_t^T = V_t^S \cup \mathcal{P}\), where \(\mathcal{P}\) are the solution positions only the teacher has.
3. Student-support projection
OPASD restricts teacher attention to the student-visible positions and renormalizes the remaining mass.
The target now lives on the student's support, but it is still solution-conditioned, because it was computed with the solution in context. The projection also preserves the teacher's relative preferences over every pair of visible positions.
4. Joint distillation
The projected teacher attention is distilled with a generalized Jensen–Shannon divergence and added to the token objective. Only the student receives gradients, and the solution is never needed at inference.
Experimental setup
- Models. Instruction-tuned Qwen3-1.7B, Qwen3-4B and Qwen3-8B, trained with LoRA on H100/H200 GPUs.
- Data. Up to 30K problem–solution pairs from the mathematical reasoning subset of OpenThoughts.
- Benchmarks. AIME 2024, AIME 2025, AIME 2026 and HMMT 2025, reported as Avg@12 accuracy.
- Baselines. The base instruction-tuned model, and token-only OPSD trained with the same data, initialization and budget.
Results
Avg@12 accuracy (%) · higher is better
OPASD is the best method at every model size and beats OPSD on every benchmark. Token-only OPSD is fragile at this scale. On Qwen3-4B it falls below the base model, while OPASD improves on it by 8.40 points.
| Method | AIME24 | AIME25 | AIME26 | HMMT25 | Average |
|---|---|---|---|---|---|
| Qwen3-8B | |||||
| Base | 60.83 | 48.89 | 51.39 | 29.17 | 47.57 |
| OPSD | 59.16 | 51.97 | 53.33 | 33.33 | 49.45 |
| OPASD | 67.50 | 57.22 | 61.11 | 35.83 | 55.42 |
| Qwen3-4B | |||||
| Base | 58.33 | 47.50 | 51.11 | 30.27 | 46.80 |
| OPSD | 51.11 | 43.88 | 46.11 | 28.33 | 42.36 |
| OPASD | 63.33 | 50.27 | 55.56 | 33.89 | 50.76 |
| Qwen3-1.7B | |||||
| Base | 33.61 | 30.28 | 31.94 | 19.17 | 28.75 |
| OPSD | 36.70 | 28.33 | 32.78 | 18.33 | 29.04 |
| OPASD | 43.05 | 35.27 | 34.44 | 23.33 | 34.02 |
Gains over OPSD are +5.97 (8B), +8.40 (4B) and +4.98 (1.7B) points. Across three training seeds on Qwen3-4B, the average is 50.88 ± 0.14.
More accurate and cheaper
The extra objective does not add cost. It removes cost. Because OPASD avoids response-length inflation, rollouts are far shorter, so the whole run is faster and lighter.

Stable training, shorter reasoning

The gain holds across budgets and sampling


A worked example

What matters
Ablations on Qwen3-1.7B · average Avg@12
| Variant | Average | |
|---|---|---|
| Training objective | ||
| Token only (OPSD) | 29.04 | </tr>|
| Attention only | 30.97 | </tr>|
| Token + attention (OPASD) | 34.02 | </tr>|
| Handling teacher-only attention | ||
| Redistribute to shared prompt positions | 31.39 | </tr>|
| Shuffled redistribution | 32.15 | </tr>|
| Student-support projection | 34.02 | </tr>|
| Attention divergence | ||
| Forward KL | 31.73 | </tr>|
| Reverse KL | 32.36 | </tr>|
| Jensen–Shannon | 32.85 | </tr>|
| Attention loss weight \(\lambda_{\mathrm{attn}}\) | ||
| 0.5 | 34.02 | </tr>|
| 1.0 | 32.85 | </tr>|
| 2.0 | 31.82 | </tr>|
| Layers distilled | ||
| Final layer only | 34.02 | </tr>|
| Final two layers | 32.54 | </tr>|
| Final three layers | 31.04 | </tr> </tbody> </table> </div>|
Citation
@article{arib2026opasd,
title = {Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning},
author = {Arib, Safaeid Hossain and Akter, Rabeya and Swapnil, Ismam Nur and
Sayeedi, Md. Faiyaz Abdullah and Mohiuddin, Tasnim and Islam, Md Mofijul},
journal = {arXiv preprint arXiv:2609.33200},
year = {2026}
}