SDPO: learn from your own correct attempts
Self-Distillation Policy Optimization (arXiv:2601.20802) samples a question several times and turns the attempts that succeed into teachers for the others. The teacher is the model itself reading the question with a correct attempt appended. The student reads the question alone. The loss pulls the student's next-token distributions toward the teacher's on the student's own attempt. The paper reports that SDPO reaches GRPO's accuracy with fewer samples.
|
Evolves |
model weights |
|
Signal |
one report per attempt with its score and its place in the sampling grid |
|
Loss family |
sdpo |
|
Package |
recipes/sdpo/ |
|
Processor |
reported feedback, one batch per sampling grid |
|
Needs |
GPUs, and a backend that captures tokens and log-probs |
|
Example |
What it does #
The harness samples each question several times through Reef and scores every attempt. It reports each attempt with its score and its place in the grid. The recipe waits for the whole grid and trains on it in one step.
How Reef implements it #
The processor is SDPOProcessor on the shared DistillProcessor (Processors). One batch is one complete grid. The teacher's request is the question with the first correct attempt by another rollout appended, in the reference implementation's words. An attempt whose question no other rollout got right keeps the plain request and a sample weight of 0. It stays in the step's mean and has no target. A grid whose attempts came from different policy versions is dropped and listed in the processor's status. A teacher sequence over max_teacher_tokens fails the step.
The sdpo loss family is a thin family on the Slime backend's distillation base (Loss families) and its defaults are the reference's. The teacher is a copy of the weights that moves 5% toward the policy after every step. The loss is the Jensen-Shannon divergence over the student's top 100 tokens plus one bucket for the rest of the vocabulary. The student picks those tokens in one forward before the step, so the trainer needs zero dropout. Each token's loss is weighted by the capped ratio between the policy and the rollout engine's log-probs.
The report contract #
A report references one attempt and carries its score and its place in the grid. teacher_context is optional feedback for the teacher and the example leaves it empty.
{
"references": ["<receipt of the attempt>"],
"score": 1.0,
"metadata": {"step": 3, "group": 12, "rollout": 5, "teacher_context": ""}
}Configuration #
groups_per_step32questions in a grid.rollouts_per_group8attempts at each question.tokenizer_pathrequiredthe served model's tokenizer directory. It renders the teacher prompt with the engine's chat template.max_teacher_prompt_tokens10240the rendered teacher prompt is cut to this many tokens.max_teacher_tokens18432a longer teacher sequence fails the step. Set it to the trainer's window.success_reward_threshold0.5an attempt at or above this score is correct.allow_own_success_as_demonstrationfalselet a correct attempt read its own response.remove_thinking_from_demonstrationtruestrip <think> blocks from the demonstration.include_environment_feedbackfalseadd the report's teacher_context to the teacher's prompt.environment_feedback_only_without_solutiontrueuse the feedback only when no correct sibling exists.enable_thinkingfalsethe chat template's thinking switch, set as the engine sampled.max_staleness0accepted lag between the producing and serving version.The Slime driver takes --loss-type custom_loss and --use-rollout-logprobs and --disable-compute-advantages-and-returns. The family adds its own flags:
--sdpo-teacherselfself is the model itself. separate is another checkpoint set by --sdpo-teacher-checkpoint.--sdpo-divergencejsdforward, reverse or jsd. --sdpo-jsd-beta is the teacher's weight in the mixture and defaults to 0.5.--sdpo-top-k100tokens per position the divergence is computed on. 0 keeps the whole distribution.--sdpo-top-k-sourcestudentwho picks the tokens, student or teacher.--sdpo-top-k-distributiontailtail keeps the rest of the vocabulary as one bucket. renormalized conditions on the picked tokens.--sdpo-teacher-update-rate0.05fraction of the policy mixed into the teacher after every step. 0 freezes the initial weights.--sdpo-importance-sampling-leveltokentoken weights each token by its own ratio. sequence averages the ratio over the response.--sdpo-importance-sampling-cap2.0cap of the importance weight. 0 disables the correction.--sdpo-skip-response-tokens0response tokens at the start of every attempt left out of the loss.Run the example #
The example trains Qwen3-8B on the Chemistry split of SciKnowEval on four GPUs. Each step samples 32 questions eight times and the test split is scored every five steps. The example's README describes the protocol.
cd recipes/sdpo/examples/sciknoweval
hf download Qwen/Qwen3-8B --local-dir ~/models/Qwen3-8B
./run.shResults #
One run of 100 steps. Test accuracy goes from 41.2% to 74.4% at step 75 and ends at 71.7%. The rollouts stay diverse and a question seen a second time is answered no better than the rest of the grid.
Related guides #
-
Inference and feedback quickstart: learn the request, receipt, and report workflow.
-
Train model weights from agent feedback: set up the GPU stack and inspect published updates.
-
Loss families: how a family such as sdpo plugs into the Slime backend, and the distillation base it is built on.