1
Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
This paper introduces label-free bias-only Test-Time Reinforcement Learning (TTRL), which uses majority-vote pseudolabels and optimizes only ~100K bias parameters. It achieves 76.67% accuracy on MATH-500 with Qwen2.5-7B, optimizing 76,000x fewer parameters than full-parameter TTRL, demonstrating substantial adaptation from a tiny subspace.
Hugging Face Daily Papersarxiv.org1 minpaper
