Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

Study Finds Focused Supervision Helps Cross-Tokenizer Distillation

Study Finds Focused Supervision Helps Cross-Tokenizer Distillation

HuggingFace·Wednesday, October 7, 2026
  • •Three teacher–student pairs tested on mathematical reasoning and code generation across different tokenizers
  • •Student-selected top-16 shared-vocabulary tokens matched full-vocabulary accuracy and beat evaluated cross-tokenizer baselines
  • •Span supervision covered all mismatched groups but reduced accuracy and produced weak or conflicting gradients
  • •Three teacher–student pairs tested on mathematical reasoning and code generation across different tokenizers
  • •Student-selected top-16 shared-vocabulary tokens matched full-vocabulary accuracy and beat evaluated cross-tokenizer baselines
  • •Span supervision covered all mismatched groups but reduced accuracy and produced weak or conflicting gradients
  • •Three teacher–student pairs tested on mathematical reasoning and code generation across different tokenizers
  • •Student-selected top-16 shared-vocabulary tokens matched full-vocabulary accuracy and beat evaluated cross-tokenizer baselines
  • •Span supervision covered all mismatched groups but reduced accuracy and produced weak or conflicting gradients
  • •Three teacher–student pairs tested on mathematical reasoning and code generation across different tokenizers
  • •Student-selected top-16 shared-vocabulary tokens matched full-vocabulary accuracy and beat evaluated cross-tokenizer baselines
  • •Span supervision covered all mismatched groups but reduced accuracy and produced weak or conflicting gradients

Researchers Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li, Wenfeng Feng, Guohua Liu and Yuewei Zhang report that on-policy distillation can work across language models with different tokenizers by prioritizing reliable supervision over broader alignment coverage. Their paper, published on October 6 and submitted to Hugging Face on October 7, tested three teacher–student pairs on mathematical reasoning and code generation. On student responses sampled before distillation, strictly aligned token positions retained nearly all teacher and student probability mass on average, despite substantial vocabulary mismatches.

At each strictly aligned position, limiting reverse KL (a measure comparing two probability distributions) to the student-selected top-16 tokens from the shared vocabulary achieved accuracy comparable to using the full shared vocabulary. It also outperformed the cross-tokenizer baselines the researchers evaluated. Strict 1:1 groups already covered most tokens generated by students, so the experiments found little benefit in expanding supervision to mismatched spans.

Adding mean squared error supervision on log probabilities across mismatched spans provided complete coverage, but reduced accuracy. At checkpoints trained only with the strict loss, gradients from span supervision had weak or negative directional agreement with strict-loss gradients and grew larger relative to them. The authors say these findings support compact supervision at strictly aligned positions over broader coverage that introduces weakly aligned or conflicting training signals.

Researchers Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li, Wenfeng Feng, Guohua Liu and Yuewei Zhang report that on-policy distillation can work across language models with different tokenizers by prioritizing reliable supervision over broader alignment coverage. Their paper, published on October 6 and submitted to Hugging Face on October 7, tested three teacher–student pairs on mathematical reasoning and code generation. On student responses sampled before distillation, strictly aligned token positions retained nearly all teacher and student probability mass on average, despite substantial vocabulary mismatches.

At each strictly aligned position, limiting reverse KL (a measure comparing two probability distributions) to the student-selected top-16 tokens from the shared vocabulary achieved accuracy comparable to using the full shared vocabulary. It also outperformed the cross-tokenizer baselines the researchers evaluated. Strict 1:1 groups already covered most tokens generated by students, so the experiments found little benefit in expanding supervision to mismatched spans.

Adding mean squared error supervision on log probabilities across mismatched spans provided complete coverage, but reduced accuracy. At checkpoints trained only with the strict loss, gradients from span supervision had weak or negative directional agreement with strict-loss gradients and grew larger relative to them. The authors say these findings support compact supervision at strictly aligned positions over broader coverage that introduces weakly aligned or conflicting training signals.

Read original (English)·Oct 7, 2026
#on policy distillation#cross tokenizer#reverse kl#shared vocabulary#token alignment#mathematical reasoning#code generation