proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYifan Yang, Xiaoyu Yang, Zengrui Jin2 min readpaperadvanced

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Summary

Pruned CTC limits CTC alignment calculations to the batch‑specific subset of tokens actually needed, eliminating the linear memory blow‑up with large vocabularies while preserving exact loss and gradient values. Applied to LLM‑based ASR, it cuts per‑step memory 5.1× with modest compute overhead and delivers near‑baseline WER at 7‑10× faster inference for both offline and bounded‑history streaming.

  • Pruned CTC restricts alignment to target tokens + blank per batch, reducing memory proportional to vocab size while keeping full‑vocab normalization.
  • The paper proves loss and gradient equivalence to standard CTC, so accuracy is unchanged.
  • Combined with finite‑beam pruning, memory per training step drops 5.1× with only ~17% extra compute time.
  • LLM‑CTC adapts pretrained LLMs (Qwen3 0.6B‑32B) for non‑autoregressive ASR, staying within 7% relative WER of cross‑entropy baselines while being 7‑10× faster.

ASR engineers working with large‑vocabulary or LLM‑based models need a memory‑efficient training objective that doesn’t sacrifice accuracy.

8/10

Related reading

  1. tokenizers v1: encode, decode and scaling, measured

    Hugging Face has released `tokenizers` v1, a major performance update that achieves 3-30x faster encoding than v0.23 while maintaining identical output and API compatibility. Key optimizations include a SIMD-accelerated splitter, a thread-local word cache, and an allocation-free BPE merge loop, ensuring tokenization doesn't bottleneck ML workflows.

    Hugging Facehuggingface.co10 min
  2. Training Compute-Optimal Large Language Models

    The paper derives a compute‑optimal scaling law showing model size and training tokens should grow together, and validates it by training a 70B‑parameter model (Chinchilla) on 1.4 T tokens that outperforms much larger LLMs.

    Hall of Famearxiv.org66 minpaper
  3. How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus

    This paper investigates the "lossless" claim of Orthrus, a hybrid architecture for accelerating LLM inference. It finds that under BF16 precision, Orthrus diverges from the exact autoregressive output trajectory in over 50% of cases, though FP32 maintains exact matching. Despite BF16 divergence, downstream task performance was not systematically degraded.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

    onPanda is an interactive annotation tool that lets humans correct LLM outputs token‑by‑token, then resumes generation from the corrected prefix. In a controlled study it cut median annotation time by 52% and the authors release a token‑level correction dataset (Panda‑CVL) for on‑policy fine‑tuning.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

    WeVisDoc introduces a two‑stage data‑centric pipeline for end‑to‑end document parsing. Stage I expands coverage using heterogeneous data and structure‑preserving degradations. Stage II probes the Stage I model with a held‑out set, clusters residual errors, and directs targeted data creation and token‑budget reallocation. The 4‑billion‑parameter model reaches 95.38 Overall on OmniDocBench v1.6 and…

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. Register Tokens for Bounded-State Reasoning in Diffusion Language Models

    Register tokens are fixed‑position embeddings that store a compact hidden state across diffusion‑based language model generation chunks, enabling bounded‑state reasoning without retaining all prior text. Post‑training on LLaDA and Dream shows up to +8.5 math and +19.5 code benchmark points versus plain text carry, and RL fine‑tuning further improves long‑horizon tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper