proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersAvinash Amballa, Yashas Malur Saidutta, Wenbo Li1 min readpaperadvanced

Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks

Summary

This paper introduces "Net Utility," a data-free metric for budget-aware LoRA merging that addresses the performance gap caused by uniform rank budget assumptions. By scoring and selecting singular directions based on task utility and interference, it achieves average performance improvements of over 2% on both vision and language tasks.

  • Existing LoRA merging methods often assume uniform rank budgets, leading to performance degradation.
  • Net Utility is a data-free metric that decomposes task LoRAs via SVD to score singular directions.
  • Each singular direction is scored based on its task utility and interference with other tasks.
  • The method globally pools these scores to select the most impactful directions under a total rank budget constraint.

Engineers optimizing LoRA merging for multi-task inference will find this relevant for improving model performance while managing resource constraints.

7/10

Related reading

  1. The Router Within: Eliciting Native Skill Routing from a Frozen LLM

    The paper introduces Gavel, a method that extracts a frozen LLM's internal routing signal via two trained linear maps, eliminating the need to embed skill descriptions in the prompt. Experiments on Qwen3‑32B show up to 13.4‑point improvements on task benchmarks and higher skill‑use accuracy compared to larger retrieval‑based systems.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. LOCI: Spatial Linear Memory for Streaming World Models

    LOCI is a hybrid spatial-memory architecture for streaming video world models, combining key-value caches with recurrent linear attention. It leverages projective camera geometry to condition memory operations, enabling more faithful reproduction of revisited content and significantly reducing peak memory usage for long videos.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

    APM-Bench is a new benchmark for evaluating persistent memory in egocentric streaming video assistants across intermittent sessions. It reveals a significant utility-latency-storage trade-off, showing current models struggle with long-term recall, low overhead, and proactive assistance simultaneously.

    Hugging Face Daily Papersarxiv.org1 minpaper