proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersMuhammad Huzaifa, Lea Schönherr, Thorsten Eisenhofer1 min readpaperadvanced

WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation

Summary

WISE-ATTA addresses budgeted active test-time adaptation (ATTA), where labels are only available for a fraction of test batches. It proposes a budget-aware strategy to decide when to request supervision over time, prioritizing useful periods and using drift-based sample selection.

  • Budgeted ATTA shifts the focus from 'what' to label within a batch to 'when' to apply supervision over time.
  • WISE-ATTA allocates supervision based on lightweight online signals, prioritizing periods where it's most useful.
  • It employs a drift-based sample selection criterion targeting samples with unconverged adaptation dynamics.
  • The method achieves competitive or improved performance on distribution shifts while requiring substantially fewer labels.

ML engineers deploying models in production environments with evolving data distributions and limited labeling budgets will find this approach valuable for maintaining model performance efficiently.

8/10

Related reading

  1. Sharpening Tax in Post-Training

    The authors introduce Sharpening Tax, a metric that measures how much RL post‑training reduces test‑time scalability (pass@K) of LLM agents. They also propose Posterior‑Tempered Group Sampling, a simple temperature‑adaptation technique that lowers this tax and improves both single‑shot accuracy and coverage.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

    This paper introduces label-free bias-only Test-Time Reinforcement Learning (TTRL), which uses majority-vote pseudolabels and optimizes only ~100K bias parameters. It achieves 76.67% accuracy on MATH-500 with Qwen2.5-7B, optimizing 76,000x fewer parameters than full-parameter TTRL, demonstrating substantial adaptation from a tiny subspace.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Mitigating the Length-Scaling Tax with Online Distillation

    The authors define the length‑scaling tax (LST) as excess response length without accuracy gain and propose Length Self‑Distillation (LSD), an online EMA‑based teacher that requires no external model. Experiments show LSD matches or exceeds RL performance while cutting LST from 19% to -3.7% on single‑turn and from 31.4% to 13.7% on multi‑turn tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Dynamically Scaled Activation Steering

    Dynamically Scaled Activation Steering (DSAS) is a method‑agnostic framework that learns per‑token, per‑layer scaling factors to turn existing activation‑steering interventions on only when a model is likely to produce undesired output (e.g., toxic text). The scaling can be optimized jointly with any steering function, improves the toxicity‑utility trade‑off on language models, transfers to text‑…

    Apple Machine Learning Researchapple.com1 minpaper
  5. onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

    onPanda is an interactive annotation tool that lets humans correct LLM outputs token‑by‑token, then resumes generation from the corrected prefix. In a controlled study it cut median annotation time by 52% and the authors release a token‑level correction dataset (Panda‑CVL) for on‑policy fine‑tuning.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

    TRACE is a condition-aware benchmark for streaming video understanding that makes evidence timing and trigger conditions explicit, and measures answer quality, timeliness, workload, and reliability. Experiments on 1,240 records show that identical QA accuracy can hide large differences in completion, false alarms, and processing cost.

    Hugging Face Daily Papersarxiv.org1 minpaper