proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJianguo Huang, Jinming Liu, Qiyao Wang1 min readpaperadvanced

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

Summary

APM-Bench is a new benchmark for evaluating persistent memory in egocentric streaming video assistants across intermittent sessions. It reveals a significant utility-latency-storage trade-off, showing current models struggle with long-term recall, low overhead, and proactive assistance simultaneously.

  • Real-world streaming assistants need persistent memory across intermittent sessions, a gap in current benchmarks.
  • APM-Bench models multi-session "life trajectories" with fine-grained video annotations and diverse questions.
  • Key challenges for persistent memory include selective retention, efficient injection, and recognizing missing evidence.
  • Evaluation reveals a clear utility-latency-storage trade-off for existing memory systems.

Researchers and engineers developing AI assistants for egocentric video will find this benchmark crucial for evaluating and improving persistent memory systems that handle real-world, intermittent interactions.

7/10

Related reading

  1. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Do LLMs Have the Memory of a Goldfish?

    The article explains that LLMs don’t have persistent personal memory; all “memory” is supplied by the surrounding application via the context window, summaries, or external storage. It outlines the distinction between trained weights, working‑memory (token context), and persistent application memory, shows how to construct API calls to preserve conversation state, and discusses the cost and laten…

    ByteByteGobytebytego.com12 min
  3. TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

    TRACE is a condition-aware benchmark for streaming video understanding that makes evidence timing and trigger conditions explicit, and measures answer quality, timeliness, workload, and reliability. Experiments on 1,240 records show that identical QA accuracy can hide large differences in completion, false alarms, and processing cost.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops

    SiliconBench benchmarks nine Apple‑Silicon LLM serving engines on unified‑memory desktops, measuring speed, memory usage, and output fidelity across Qwen3, Qwen3.5, and Gemma‑4 models. It finds vllm‑metal doubles throughput at modest concurrency, memory budgets often fail to preserve headroom, and only three stacks satisfy all fidelity and model‑coverage requirements, with tensor‑parallel scaling…

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Shared Selective Persistent Memory for Agentic LLM Systems

    Apple proposes a memory architecture for agentic LLMs that selectively persists reusable context (specs, schemas, configs, constraints) across sessions and users. Shared workspaces with role‑based access and a zero‑token data‑refresh mechanism cut token usage by 97×, reduce task time by 14×, and raise task‑completion rates to 96% versus 71%‑79% for baselines.

    Apple Machine Learning Researchapple.com1 minpaper