Skip to content
Research noteAI-2026-0223

A budget for how much the teacher says

Self-distillation always aims at the full teacher. Constraining how much teacher information enters at each token improved specialisation in seven of eight settings.

3 minMain AI Hub

Teaching a model a new skill without damaging the ones it already has is the practical problem behind most fine-tuning work. Self-distillation fine-tuning addresses it by having the model learn from a version of itself conditioned on demonstrations, which reduces forgetting compared with training on the demonstrations directly.

A paper posted this week argues that the method contains an unexamined assumption: it always distils toward the full demonstration-conditioned teacher. Teacher influence is fixed at that endpoint, and there is no control over how much of the demonstration should be transferred at any particular point in the sequence.

The proposed alternative treats the teacher as a budgeted source rather than a target. At each token the method selects the distribution closest to the current student that still satisfies a prescribed constraint on how much teacher information may enter — which yields a closed-form target with a tilt determined locally rather than globally. To stop cumulative drift, the student is additionally anchored to its own frozen base policy.

Across four different model backbones and two specialisation tasks, the method improved on ordinary self-distillation in seven of eight model-task combinations and matched it in the eighth. On retention the gap is larger: 73% of evaluations stayed within half a point of the base model, against 52% for the strongest baseline.

It also produced the largest mean improvement across ten additional benchmarks in mathematics, coding and competition mathematics — which is the combination that matters. Retention and specialisation usually trade against each other, and a method that improves both is describing something structural rather than a tuning win.

The finding underneath is that when teacher information arrives matters as much as how much of it does. That is a knob nobody was turning, because the standard formulation does not expose one.

Retold from arXiv. This is a summary in our own words; follow the link for the original reporting.

Read next

Across the network

Desks that share a zone with this one on the BITBRIEF coverage map.

Terms defined