arXiv 2026 New LLM ReasoningPost-TrainingDistillation

Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

Safaeid Hossain Arib1, Rabeya Akter2, Ismam Nur Swapnil1, Md. Faiyaz Abdullah Sayeedi3, Tasnim Mohiuddin4,*, Md Mofijul Islam5,*,‡

1ACI PLC, Bangladesh   2University of Dhaka   3BRAC University   4QCRI, Qatar   5Amazon GenAI, USA

arXiv preprint arXiv:2609.33200, 2026

* Equal supervision  ·  ‡ Work done outside role at Amazon

TL;DR

A privileged teacher that has seen the verified solution knows not only what token comes next but also which earlier steps matter. OPASD distills that attention too, after projecting it onto the context the student can actually see, and gets more accurate, shorter, and cheaper reasoning.

  • +8.40pt Avg@12 over OPSD on Qwen3-4B (+4.98 to +8.40 across 3 sizes)
  • −73.9%generated rollout tokens vs. token-only OPSD
  • −72.6%estimated training compute (EFLOPs)
  • 1.53×faster wall-clock training
Overview figure for Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning
OPSD vs. OPASD. (a) OPSD aligns teacher and student next-token distributions, telling the student what to predict. (b) OPASD additionally aligns projected attention distributions, telling it where to look.

Abstract

On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53× faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.

What to predict is not where to look

Motivation

Reinforcement learning with verifiable rewards gives a reasoning model one number per response. On-policy distillation is denser: a teacher scores every prefix of the student's own trajectory with a full next-token distribution. On-policy self-distillation (OPSD) goes further and needs no separate teacher. The same model plays both roles, and the "teacher" copy is simply shown the verified solution.

That privileged teacher knows more than the next token. Multi-step reasoning hinges on constraints and intermediate results established much earlier, and the teacher's attention shows which of them matter at each step. OPSD discards this signal. OPASD distills it.

The catch is a support mismatch

Next-token distributions are directly comparable because teacher and student share one vocabulary. Attention distributions are not. They live over context positions, and the teacher's context contains solution tokens that the student never sees. Some of the teacher's attention mass sits on positions that do not exist for the student, so the two distributions cannot be aligned as they are.

Method

On-policy sampling, two views, projection, joint distillation

OPASD method overview with on-policy sampling, teacher-student distributions, student-support projection, and joint distillation
OPASD overview. (1) The student samples a trajectory from the problem alone. (2) Student and privileged teacher evaluate the same trajectory, producing token and attention distributions. (3) Teacher attention on solution-only positions is removed and the rest renormalized. (4) Token and projected attention distributions are distilled jointly.

1. On-policy self-distillation

One language model defines two policies. The teacher is the frozen initial checkpoint conditioned on the problem and its verified solution, and the student conditions on the problem alone, exactly as at inference time.

\[ p_T(\cdot \mid x, y^\star) \triangleq p_{\theta_0}(\cdot \mid x, y^\star), \qquad p_S(\cdot \mid x) \triangleq p_\theta(\cdot \mid x) \]

The student samples its own trajectory \(\hat{y} \sim p_S(\cdot \mid x)\), and the teacher scores that same trajectory token by token rather than generating its own. The token-level objective matches their next-token distributions at every step.

\[ \mathcal{L}_{\mathrm{tok}} = \frac{1}{T} \sum_{t=1}^{T} D_{\mathrm{tok}}\big(p_t^T, p_t^S\big), \qquad p_t^S = p_S(\cdot \mid x, \hat{y}_{<t}), \;\; p_t^T = p_T(\cdot \mid x, y^\star, \hat{y}_{<t}) \]

2. Attention from both views

For each generated token, OPASD reads the head-averaged final-layer attention of both models. The student attends over its visible support \(V_t^S\) (problem and prefix). The teacher attends over \(V_t^T = V_t^S \cup \mathcal{P}\), where \(\mathcal{P}\) are the solution positions only the teacher has.

3. Student-support projection

OPASD restricts teacher attention to the student-visible positions and renormalizes the remaining mass.

\[ \widetilde{A}_t^T(i) = \frac{A_t^T(i)}{\sum_{j \in V_t^S} A_t^T(j)}, \qquad i \in V_t^S \]

The target now lives on the student's support, but it is still solution-conditioned, because it was computed with the solution in context. The projection also preserves the teacher's relative preferences over every pair of visible positions.

\[ \frac{\widetilde{A}_t^T(i)}{\widetilde{A}_t^T(j)} = \frac{A_t^T(i)}{A_t^T(j)} \qquad \forall\, i, j \in V_t^S \]

4. Joint distillation

The projected teacher attention is distilled with a generalized Jensen–Shannon divergence and added to the token objective. Only the student receives gradients, and the solution is never needed at inference.

\[ \ell_{\mathrm{attn}}(t) = \beta\, D_{\mathrm{KL}}\big(\widetilde{A}_t^T \,\Vert\, M_t\big) + (1-\beta)\, D_{\mathrm{KL}}\big(A_t^S \,\Vert\, M_t\big), \qquad M_t = \beta \widetilde{A}_t^T + (1-\beta) A_t^S \]
\[ \mathcal{L} = \mathcal{L}_{\mathrm{tok}} + \lambda_{\mathrm{attn}}\, \mathcal{L}_{\mathrm{attn}} \]

Experimental setup

  • Models. Instruction-tuned Qwen3-1.7B, Qwen3-4B and Qwen3-8B, trained with LoRA on H100/H200 GPUs.
  • Data. Up to 30K problem–solution pairs from the mathematical reasoning subset of OpenThoughts.
  • Benchmarks. AIME 2024, AIME 2025, AIME 2026 and HMMT 2025, reported as Avg@12 accuracy.
  • Baselines. The base instruction-tuned model, and token-only OPSD trained with the same data, initialization and budget.

Results

Avg@12 accuracy (%) · higher is better

OPASD is the best method at every model size and beats OPSD on every benchmark. Token-only OPSD is fragile at this scale. On Qwen3-4B it falls below the base model, while OPASD improves on it by 8.40 points.

MethodAIME24AIME25AIME26HMMT25Average
Qwen3-8B
Base60.8348.8951.3929.1747.57
OPSD59.1651.9753.3333.3349.45
OPASD67.5057.2261.1135.8355.42
Qwen3-4B
Base58.3347.5051.1130.2746.80
OPSD51.1143.8846.1128.3342.36
OPASD63.3350.2755.5633.8950.76
Qwen3-1.7B
Base33.6130.2831.9419.1728.75
OPSD36.7028.3332.7818.3329.04
OPASD43.0535.2734.4423.3334.02

Gains over OPSD are +5.97 (8B), +8.40 (4B) and +4.98 (1.7B) points. Across three training seeds on Qwen3-4B, the average is 50.88 ± 0.14.

More accurate and cheaper

The extra objective does not add cost. It removes cost. Because OPASD avoids response-length inflation, rollouts are far shorter, so the whole run is faster and lighter.

Accuracy versus wall-clock time for OPASD and OPSD
OPASD Pareto-dominates OPSD on Qwen3-4B. Higher Avg@12 in 5.27 h instead of 8.04 h, with 15.58M instead of 59.73M generated tokens, 0.67 instead of 2.43 EFLOPs, and 11.9% less peak GPU memory.

Stable training, shorter reasoning

Training and validation accuracy, response length, and epistemic token usage for OPSD and OPASD
Training dynamics on Qwen3-4B. OPSD's accuracy drops while its responses balloon and it leans heavily on "wait", "maybe" and "alternatively". OPASD stays stable with concise responses and small shifts in epistemic-token usage.

The gain holds across budgets and sampling

Avg@12 across 16K, 28K and 38K test-time generation budgets
Test-time generation budget. Trained with an 8K response cap and evaluated at 16K, 28K and 38K tokens, OPASD beats OPSD on every benchmark at every budget, by 8.40, 6.25 and 5.81 points on average.
Pass@1 and Pass@12 over training steps for OPSD and OPASD
Pass@1 and Pass@12 over training on Qwen3-4B. OPASD achieves higher Pass@1 on all four benchmarks and keeps an overall advantage at Pass@12.

A worked example

Reasoning traces of the base model, OPSD and OPASD on an AIME24 number theory problem
AIME24 number theory. The base model overlooks valid alternatives. OPSD considers more candidates but evaluates one incorrectly. OPASD carries the intermediate result through and finds the correct minimum, 110.

What matters

Ablations on Qwen3-1.7B · average Avg@12

</tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tr> </tbody> </table> </div>
  • The two signals are complementary. Attention alone already beats token-only OPSD, and combining them adds another 3 points.
  • Projection is the right fix. Moving the teacher-only mass elsewhere, even randomly, is worse than simply removing it and renormalizing.
  • Light, late supervision works best. A moderate weight on the final layer only gives the highest accuracy.

Limitations

OPASD is evaluated on three Qwen3 model sizes, and testing other model families would give a more complete picture. The privileged teacher is kept fixed at the initial checkpoint, and whether attention supervision helps as much under other teacher-update strategies remains open.

VariantAverage
Training objective
Token only (OPSD)29.04
Attention only30.97
Token + attention (OPASD)34.02
Handling teacher-only attention
Redistribute to shared prompt positions31.39
Shuffled redistribution32.15
Student-support projection34.02
Attention divergence
Forward KL31.73
Reverse KL32.36
Jensen–Shannon32.85
Attention loss weight \(\lambda_{\mathrm{attn}}\)
0.534.02
1.032.85
2.031.82
Layers distilled
Final layer only34.02
Final two layers32.54
Final three layers31.04

Citation

@article{arib2026opasd,
  title   = {Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning},
  author  = {Arib, Safaeid Hossain and Akter, Rabeya and Swapnil, Ismam Nur and
             Sayeedi, Md. Faiyaz Abdullah and Mohiuddin, Tasnim and Islam, Md Mofijul},
  journal = {arXiv preprint arXiv:2609.33200},
  year    = {2026}
}