Hugging Face Daily PapersSuvradip Paul, Chandra Bhushan, Harsh Sharma1 min readpaperadvanced
IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
Summary
IndicBankBench is a 799‑case benchmark for Indian retail banking assistants that evaluates safety, tool use, response adequacy, and advisory quality across multiple trials. The study finds strict reliability of current LLM assistants is only 44‑58%, exposing a sizable gap between apparent and dependable performance.
- 799 test cases span five banking domains and 20 primary axes, evaluated in four stages: safety, tool use, response adequacy, advisory quality.
- Deterministic checks cover tool use and safety; an LLM judge assesses semantic adequacy, with each case run three times to compute strict‑pass and at‑least‑once metrics.
- Across 11 evaluated models, strict reliability ranges 43.7‑58.2% while at‑least‑once success is 60‑74%, highlighting over‑optimistic success reporting.
- The benchmark pinpoints failure modes such as unnecessary clarification, stale context, wrong account selection, or mismatched value writes.
Financial product teams and safety engineers building LLM assistants should care because the benchmark reveals hidden reliability gaps in real‑world banking interactions.
7/10
