proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYibo Ma, Qianqian Zhang, Peng Liu1 min readpaperadvanced

TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

Summary

TRACE is a condition-aware benchmark for streaming video understanding that makes evidence timing and trigger conditions explicit, and measures answer quality, timeliness, workload, and reliability. Experiments on 1,240 records show that identical QA accuracy can hide large differences in completion, false alarms, and processing cost.

  • TRACE adds evidence‑timing and instruction‑dependent trigger annotations to streaming video tasks.
  • The Core‑Adapter protocol controls information flow while logging history processing and response events.
  • Metrics go beyond accuracy: they include timeliness, workload, false‑alarm rate, missed windows, and reliability.
  • Eight public models achieve similar QA scores but differ widely in completion rates and generation cost.

Anyone building or researching streaming video QA systems should care, because single‑score evaluations miss critical operational trade‑offs.

7/10

Related reading

  1. APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

    APM-Bench is a new benchmark for evaluating persistent memory in egocentric streaming video assistants across intermittent sessions. It reveals a significant utility-latency-storage trade-off, showing current models struggle with long-term recall, low overhead, and proactive assistance simultaneously.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning

    E2A‑Bench is a 969‑query benchmark for financial chart reasoning that evaluates vision‑language models across a full evidence‑to‑action chain using four metrics (UCR, RCI, ECI, NDR). Experiments on 20 VLMs expose hidden failures: low‑UCR models have only 6.4 % directional coverage, oracle‑aided verification cuts unsupported claims but can kill coverage, and fine‑tuning inflates BUY:SELL ratios by…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Dapper, a Large-Scale Distributed Systems Tracing Infrastructure

    Dapper is Google’s production‑grade distributed tracing system that achieves low overhead and ubiquitous deployment by instrumenting only a few core libraries and using adaptive sampling. The paper details its data model, sampling strategy, and the ecosystem of analysis tools built on top of it.

    Hall of Famegoogle.com43 minpaper