proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersSuvradip Paul, Chandra Bhushan, Harsh Sharma1 min readpaperadvanced

IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking

Summary

IndicBankBench is a 799‑case benchmark for Indian retail banking assistants that evaluates safety, tool use, response adequacy, and advisory quality across multiple trials. The study finds strict reliability of current LLM assistants is only 44‑58%, exposing a sizable gap between apparent and dependable performance.

  • 799 test cases span five banking domains and 20 primary axes, evaluated in four stages: safety, tool use, response adequacy, advisory quality.
  • Deterministic checks cover tool use and safety; an LLM judge assesses semantic adequacy, with each case run three times to compute strict‑pass and at‑least‑once metrics.
  • Across 11 evaluated models, strict reliability ranges 43.7‑58.2% while at‑least‑once success is 60‑74%, highlighting over‑optimistic success reporting.
  • The benchmark pinpoints failure modes such as unnecessary clarification, stale context, wrong account selection, or mismatched value writes.

Financial product teams and safety engineers building LLM assistants should care because the benchmark reveals hidden reliability gaps in real‑world banking interactions.

7/10

Related reading

  1. E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning

    E2A‑Bench is a 969‑query benchmark for financial chart reasoning that evaluates vision‑language models across a full evidence‑to‑action chain using four metrics (UCR, RCI, ECI, NDR). Experiments on 20 VLMs expose hidden failures: low‑UCR models have only 6.4 % directional coverage, oracle‑aided verification cuts unsupported claims but can kill coverage, and fine‑tuning inflates BUY:SELL ratios by…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. ROWBench: Do Video Models Render What the Program Specifies?

    PROWBench is a new benchmark designed to evaluate the visual fidelity of programmable world models to fine-grained, program-specified events and interactions. It comprises 170 programmatically constructed episodes and 600 proxy videos, using VLM-based metrics to check generated videos against observable consequences of program execution.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

    PACT is a new benchmark designed to measure how well enterprise LLM agents follow compliance rules, especially when under user pressure. It found substantial variability across 22 models, with even the strongest assistants misapplying rules 6-10% of the time, and user pressure increasing violation rates by 65% on average.

    Hugging Face Daily Papersarxiv.org2 minpaper
  4. How UK AISI and EvalEval Are Making Benchmark Results Reproducible

    The UK AI Security Institute (AISI) is publishing its frontier‑LLM benchmark results on EvalEval’s open Evaluation Cards platform, using the Every Eval Ever (EEE) schema. The release covers five main benchmarks (HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, Terminal‑Bench 2.0) and six models (Claude Opus 4/4.5/4.6, GPT‑5/5.2/5.4), plus two cyber‑evaluation suites. The data inclu…

    Hugging Facehuggingface.co3 min
  5. Constraint Decay: The Fragility of LLM Agents in Backend Code Generation

    A systematic evaluation of LLM agents generating multi‑file backend code shows a sharp drop in correctness when structural constraints (framework conventions, ORM usage, API contracts) are added. Across 100 tasks in 8 Python web frameworks, assertion pass rates fall ~27 points, with data‑layer bugs (bad queries, ORM violations) driving most failures. Mid‑size models cope with minimal frameworks (…

    arXiv cs.SE (Software Engineering)arxiv.org1 minpaperHN287197
  6. Article: Architecting Secure and Scalable Facial Verification Systems

    A real‑world post‑mortem of a high‑volume face verification service that moved from a naïve synchronous API to an async, layered pipeline (edge validation, preprocessing, decoupled detection/verification, decision engine) to achieve 8.5k rpm, p99 < 1.8 s, 30 % cost savings, and strict privacy controls.

    InfoQinfoq.com15 min