1
Calibrating Teacher--Student Discrepancy for On-Policy Distillation
This paper introduces Calibrated On-Policy Distillation (Cal-OPD), a method to improve knowledge distillation by filtering out the teacher model's inherent deviations from the learning signal. Cal-OPD estimates and removes the teacher's self-deviation, leading to better student performance on mathematical reasoning tasks.
Hugging Face Daily Papersarxiv.org1 minpaper
