proomt

Search

Search posts, papers, and topics

model training

RSS
  1. 2

    On the Off-Policy Teacher in On-Policy Distillation

    On-policy distillation (OPD) faces a challenge where the teacher's performance degrades when supervising student-generated, off-policy trajectories. The SCOUT framework addresses this by co-training the teacher to adapt to student-generated prefixes using reinforcement learning, consistently improving OPD effectiveness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. 3

    From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

    This paper introduces SBERT2S1, a method to convert biomedical sentence encoders into typed decision models, and BIODECIDE, a new biomedical typed-decision suite. It finds that retrieval training benefits prior-fused residual (PFR) models, cross-head (C) models generally outperform PFR, and the RLCD objective trails cross-entropy due to reward normalization issues.

    Hugging Face Daily Papersarxiv.org1 minpaper