proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersNaveen Vakada, Mingyuan Li, Shaoxiong Ji1 min readpaperadvanced

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

Summary

This paper introduces label-free bias-only Test-Time Reinforcement Learning (TTRL), which uses majority-vote pseudolabels and optimizes only ~100K bias parameters. It achieves 76.67% accuracy on MATH-500 with Qwen2.5-7B, optimizing 76,000x fewer parameters than full-parameter TTRL, demonstrating substantial adaptation from a tiny subspace.

  • TTRL can be effective by optimizing only ~100K bias parameters, keeping the pretrained backbone frozen.
  • Majority-vote pseudolabels provide sufficient reward signals for label-free test-time adaptation.
  • The method achieved 76.67% accuracy on MATH-500 with Qwen2.5-7B, comparable to labeled bias-steering.
  • It optimizes 76,000x fewer parameters than full-parameter TTRL, a significant efficiency gain.

This work is significant for LLM researchers and practitioners seeking highly efficient, label-free methods for improving model reasoning and adaptation at inference time.

8/10

Related reading

  1. Harness-Zero: Harness Distillation via Agent-as-Harness

    Harness-Zero proposes a harness‑distillation technique where a specialized harness guides a student model via an intermediate agent‑as‑harness, allowing the learned behavior to be baked into the model weights and removed at deployment. Experiments show macro‑average task success jumps from 23.3% to 44.3%, surpassing the 41.7% achieved with the harness still attached, and recovers 82.3% of harness…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

    RetireOPD introduces a self‑retiring on‑policy distillation framework for multi‑turn RL agents. A skill‑conditioned teacher is first trained with environment rewards, then a skill‑free student learns jointly via RL and token‑level distillation. The student automatically drops the teacher once its performance gap stops shrinking and it reaches a target success‑rate fraction, after which training c…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. DataFlex-RL: An Evaluation Platform for RLVR Data Policies

    The paper introduces DataFlex‑RL, a platform to benchmark how different data‑selection policies affect reinforcement‑learning‑with‑verifiable‑rewards training. Across extensive experiments on Qwen2.5‑7B and Llama‑3.1‑8B, uniform sampling is the only method that consistently improves performance, and no alternative policy yields a statistically significant gain.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

    PLC‑DPO extends Direct Preference Optimization by using the policy‑reference margin to route each training pair into clean, flipped, or tie categories, actively correcting noisy or ambiguous labels. Across extensive benchmarks it improves mean win‑rate from 55.5 % to 60.5 % and stays stable under injected noise and tie stress tests.

    Hugging Face Daily Papersarxiv.org1 minpaper