Hugging Face Daily PapersYifan Yang, Xiaoyu Yang, Zengrui Jin2 min readpaperadvanced
Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Summary
Pruned CTC limits CTC alignment calculations to the batch‑specific subset of tokens actually needed, eliminating the linear memory blow‑up with large vocabularies while preserving exact loss and gradient values. Applied to LLM‑based ASR, it cuts per‑step memory 5.1× with modest compute overhead and delivers near‑baseline WER at 7‑10× faster inference for both offline and bounded‑history streaming.
- Pruned CTC restricts alignment to target tokens + blank per batch, reducing memory proportional to vocab size while keeping full‑vocab normalization.
- The paper proves loss and gradient equivalence to standard CTC, so accuracy is unchanged.
- Combined with finite‑beam pruning, memory per training step drops 5.1× with only ~17% extra compute time.
- LLM‑CTC adapts pretrained LLMs (Qwen3 0.6B‑32B) for non‑autoregressive ASR, staying within 7% relative WER of cross‑entropy baselines while being 7‑10× faster.
ASR engineers working with large‑vocabulary or LLM‑based models need a memory‑efficient training objective that doesn’t sacrifice accuracy.
8/10