proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersQiaolin Wang, Pedro Sandoval-Segura, Anunaya Joshi1 min readpaperadvanced

NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech

Summary

NVAlign introduces a post‑training direct‑gradient framework that uses an NV‑aware ASR model as a reward to improve non‑verbal vocalization tag following in continuous autoregressive flow‑matching TTS. Experiments on NVV‑SuperBench and human listening show higher tag‑following accuracy than fine‑tuning and Flow‑GRPO baselines.

  • NVAlign freezes a NV‑ASR model as a reward function and backpropagates through the flow‑matching sampler to update the TTS backbone.
  • A two‑step gradient surrogate enables efficient reward gradient computation despite the stochastic sampler.
  • Fidelity penalties and reference‑velocity regularization preserve speaker identity and overall speech quality during post‑training.
  • Empirical results show NVAlign outperforms supervised fine‑tuning and Flow‑GRPO on tag‑following accuracy.

TTS engineers who need precise control over non‑verbal cues (e.g., laughter, breaths) can adopt NVAlign to improve tag fidelity without retraining the entire model.

6/10

Related reading

  1. SteerDuplex: Steerable Duplex Speech Dialogue Models

    The paper presents SteerDuplex, a full‑duplex speech dialogue model that can be steered along tone, persona, and speed via instruction following, and introduces the SteerBench benchmark to evaluate such steerability. Supervised training yields a 44.5 % pass‑rate lift, and reinforcement‑learning fine‑tuning improves interruption handling and reduces pause barge‑ins, though reward hacking remains a…

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation

    Apple researchers propose Trajectory‑Shaped Discrete Flow Matching (TS‑DFM), a training‑time distillation method that replaces blind stochastic jumps in discrete flow‑matching with an energy‑based compass to select higher‑quality intermediate tokens. On a 170 M‑parameter language model, the 8‑step student outperforms the 1 024‑step teacher by 32 % perplexity while being 128× faster, beating basel…

    Apple Machine Learning Researchapple.com1 minpaper
  3. Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

    Researchers train a compact 82 M‑parameter Thai fixed‑voice TTS model using synthetic speech generated by a large voice‑cloning teacher, requiring only a 15‑second real reference. The student achieves 68.2% keyword accuracy and 91.4% pause precision, outperforming its teacher on pause placement and enabling on‑device inference.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

    The paper presents AV‑GRPO, a diffusion‑based reinforcement‑learning framework that treats audio and video generation as separate but coordinated tasks by anchoring rollouts to each modality and freezing the opposite tower during optimization. Evaluated on the new 5DAV dataset and benchmarks, it achieves higher fidelity, better text‑modality alignment, and tighter audio‑video sync than the previo…

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control

    Zing‑0.5 is a 5 B autoregressive world model that lets users control generated environments in real time using both keyboard actions and text prompts. The paper introduces unified action‑text conditioning, segment‑level teacher distillation, and a low‑cost streaming inference pipeline that runs at 24 FPS (832×480) for about $0.009 per minute, achieving 81 % overall and 88.5 % consistency on a nav…

    Hugging Face Daily Papersarxiv.org1 minpaper