1
NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech
NVAlign introduces a post‑training direct‑gradient framework that uses an NV‑aware ASR model as a reward to improve non‑verbal vocalization tag following in continuous autoregressive flow‑matching TTS. Experiments on NVV‑SuperBench and human listening show higher tag‑following accuracy than fine‑tuning and Flow‑GRPO baselines.
Hugging Face Daily Papersarxiv.org1 minpaper
