1
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
Researchers train a compact 82 M‑parameter Thai fixed‑voice TTS model using synthetic speech generated by a large voice‑cloning teacher, requiring only a 15‑second real reference. The student achieves 68.2% keyword accuracy and 91.4% pause precision, outperforming its teacher on pause placement and enabling on‑device inference.
Hugging Face Daily Papersarxiv.org1 minpaper
