proomt

Search

Search posts, papers, and topics

All posts

Hacker News front pageFrancis Bach22 min readadvanced

Exploding variance of means of exponentials: least-squares to the rescue

Summary

This article addresses the exploding variance problem when estimating log-sum-exp functions, common in machine learning, using empirical averages. It proposes a novel approach that re-frames the log-sum-exp as an integral over weighted chi-square divergences, allowing for stable, closed-form least-squares estimators.

  • Estimating log-sum-exp functions via empirical averages suffers from exponentially exploding variance with increasing potential function values.
  • Least-squares methods offer computational and statistical simplicity but are traditionally ill-suited for log-sum-exp problems.
  • The author proposes expressing KL divergence (and thus log-sum-exp) as an integral over weighted chi-square divergences.
  • Each weighted chi-square divergence has a quadratic variational form solvable by least-squares, leading to more stable estimators.

This work offers a novel, more stable method for estimating fundamental quantities in machine learning, such as log-partition functions and KL divergences, which are prone to high variance with standard sampling.

8/10

Related reading

  1. VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

    VC-Attention introduces a training‑free low‑bit attention pipeline for diffusion transformers. It smooths value tensors via lightweight online clustering (V‑Smooth) and quantizes only the residual after subtracting block means, restoring the mean from the softmax row sum. It also replaces the FP32 softmax exponential with a fused FP8 cast (ExpCast‑FP8) that maps log‑scores directly to E4M3 probab…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Learning to solve hard problems in RL for LLMs by never giving up

    The post introduces the *Matthew Effect* in RL‑fine‑tuning of LLMs—performance gains concentrate on tasks the model already solves— and proposes *Never Give Up* (NGU), an adaptive sampling scheme that uses a small k for easy prompts and retries hard prompts with a high‑probability “never give up” loop. Experiments on math (AIME, GSM8k), code (Manufactoria), and larger‑scale setups (DeepScaler) sh…

    Hacker News front pagegithub.io11 minHN1179
  3. Scaling Laws for Looped Mixture of Experts

    This paper introduces Loop Scaling Laws, which jointly model recurrence and sparsity in Mixture-of-Experts (MoE) transformers. It finds that recurrence and sparsity offer complementary efficiency gains, improving prediction accuracy and enabling more efficient large model design.

    Hugging Face Daily Papersarxiv.org1 minpaperHN2
  4. Hidden Technical Debt in Machine Learning Systems

    This paper introduces the concept of technical debt in machine learning systems, arguing that ML systems accrue unique and significant maintenance costs beyond traditional software engineering. It identifies several ML-specific risk factors like entanglement, hidden feedback loops, and data dependencies that erode system boundaries and increase long-term operational expenses.

    Hall of Famenips.cc25 minpaper
  5. 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

    The paper studies gradient‑estimation noise in sparse on‑policy distillation (OPD) and introduces the Information‑Efficiency Ratio (IER), a signal‑to‑noise based metric for selecting which tokens to supervise. IER is derived from an information‑geometry analysis with an optimal scalar baseline, and a candidate‑set approximation lets it be combined with existing usefulness scores while keeping the…

    Hugging Face Daily Papersarxiv.org1 minpaper