Hugging Face Daily PapersNaveen Vakada, Mingyuan Li, Shaoxiong Ji1 min readpaperadvanced
Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
Summary
This paper introduces label-free bias-only Test-Time Reinforcement Learning (TTRL), which uses majority-vote pseudolabels and optimizes only ~100K bias parameters. It achieves 76.67% accuracy on MATH-500 with Qwen2.5-7B, optimizing 76,000x fewer parameters than full-parameter TTRL, demonstrating substantial adaptation from a tiny subspace.
- TTRL can be effective by optimizing only ~100K bias parameters, keeping the pretrained backbone frozen.
- Majority-vote pseudolabels provide sufficient reward signals for label-free test-time adaptation.
- The method achieved 76.67% accuracy on MATH-500 with Qwen2.5-7B, comparable to labeled bias-steering.
- It optimizes 76,000x fewer parameters than full-parameter TTRL, a significant efficiency gain.
This work is significant for LLM researchers and practitioners seeking highly efficient, label-free methods for improving model reasoning and adaptation at inference time.
8/10