Hugging Face Daily PapersQiaolin Wang, Pedro Sandoval-Segura, Anunaya Joshi1 min readpaperadvanced
NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech
Summary
NVAlign introduces a post‑training direct‑gradient framework that uses an NV‑aware ASR model as a reward to improve non‑verbal vocalization tag following in continuous autoregressive flow‑matching TTS. Experiments on NVV‑SuperBench and human listening show higher tag‑following accuracy than fine‑tuning and Flow‑GRPO baselines.
- NVAlign freezes a NV‑ASR model as a reward function and backpropagates through the flow‑matching sampler to update the TTS backbone.
- A two‑step gradient surrogate enables efficient reward gradient computation despite the stochastic sampler.
- Fidelity penalties and reference‑velocity regularization preserve speaker identity and overall speech quality during post‑training.
- Empirical results show NVAlign outperforms supervised fine‑tuning and Flow‑GRPO on tag‑following accuracy.
TTS engineers who need precise control over non‑verbal cues (e.g., laughter, breaths) can adopt NVAlign to improve tag fidelity without retraining the entire model.
6/10
