Hugging Face Daily PapersBhavana Akkiraju, Ravi Sastry Kolluru, Sri Charan D1 min readpaperadvanced
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
Summary
The authors present VākQA, a 2,001‑question spoken factoid QA benchmark for Telugu with audio, transcriptions, and human‑verified answers, and they validate automatic evaluation methods against human ratings. Using this setup they show that translation loses cultural nuance, ASR errors alter meaning, and cascaded ASR‑MT errors degrade model performance.
- VākQA offers 2,001 Telugu factoid QA pairs (2.53 h audio) with bilingual transcriptions and human‑verified reference answers.
- Human studies find Gemini-as-a-judge best matches human scores but is stricter; open‑weight judges penalize correct answers with different surface forms.
- Experiments reveal translation erases cultural specificity, speech input introduces phonetic confusions, and ASR‑MT error cascades compound performance loss.
- Both proprietary and open‑weight models underperform humans, especially on spoken input and domain‑specific phrasing.
Anyone building multilingual spoken QA systems needs a realistic Telugu benchmark and reliable evaluation methods to gauge real‑world performance.
7/10