1
Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation
The paper shows that when specialist LLMs are trained only on QA pairs (no explicit reasoning supervision), their optimization implicitly selects a latent distribution of reasoning trajectories. By treating the distilled student as an agnostic probe—since it inherits only the sampled trajectories—the authors empirically demonstrate a strong correlation (across 27 specialist‑student pairs) between…
Hugging Face Daily Papersarxiv.org1 minpaper
