proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersKeyu Wang, Yangyi Huang, Jiale Kang1 min readpaperadvanced

DepthBench: Measuring How Residual Connections Enable More Computational Depth

Summary

DepthBench is a new benchmark designed to measure how effectively Transformer architectures utilize increased computational depth. The study found that residual connection designs, particularly Highway Connections (HC) and Full AttnRes, are crucial for translating architectural depth into effective computational gains, unlike standard normalization methods.

  • Deeper Transformers don't always yield more effective computation; the benefit is strongly architecture-dependent.
  • Standard Pre-LN and most norm/scaling variants show little benefit or degrade performance in deep, narrow models.
  • Highway Connections (HC) and Full AttnRes consistently improve performance with increased depth, even at extreme deep shapes.
  • Gains from HC and Full AttnRes extend beyond pre-training loss to improved domain-specific performance.

This paper provides critical insights for ML engineers and researchers designing large Transformer models, showing how to effectively scale models by depth rather than just width.

8/10

Related reading

  1. Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

    The authors cast transformer block removal as a constrained binary optimization problem equivalent to an Ising glass, using a Hessian‑derived energy as a proxy for downstream quality. Solving the resulting QUBO with classical or quantum‑inspired solvers yields up to 23 MMLU points improvement over prior block‑removal baselines at 50 % depth compression.

    Hugging Facehuggingface.co8 minHN2
  2. ROWBench: Do Video Models Render What the Program Specifies?

    PROWBench is a new benchmark designed to evaluate the visual fidelity of programmable world models to fine-grained, program-specified events and interactions. It comprises 170 programmatically constructed episodes and 600 proxy videos, using VLM-based metrics to check generated videos against observable consequences of program execution.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing

    WhiteMatter introduces all-to-all cross-layer connections in Transformers by mixing past-token representations from any depth into shared KV cache channels. This approach allows for performance comparable to 50% larger standard Transformers or improved performance with half the KV cache size. It also addresses training slowdowns with a novel cyclic iteration method.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

    VA‑Bench is a new benchmark that evaluates general‑purpose multimodal LLMs on the full observe‑reason‑act‑revise loop in embodied robotics, using RGB demonstrations, active camera control, and metric Cartesian commands. The best model reaches 53.9% average task success, showing active perception helps but long‑horizon tasks remain unsolved.

    Hugging Face Daily Papersarxiv.org1 minpaper