proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersChaoqian Ouyang, Ling Yue, Libin Zheng1 min readpaperadvanced

TokenCast: Forecasting Token Consumption During LLM Agent Execution

Summary

LLM agent token consumption varies widely and is hard to predict due to dynamic execution and growing context. TokenCast proposes a composable cost representation for execution segments to forecast token usage, reducing prediction error and enabling better budget control.

  • LLM agent token consumption can vary by over an order of magnitude for the same task.
  • TokenCast learns a composable cost representation for each execution segment, including its own consumption and context growth.
  • It accounts for cumulative context growth, which inflates input size for subsequent LLM calls.
  • TokenCast reduces mean absolute error by 14.5% against strong comparators across various tasks and models.

Engineers deploying LLM agents can use this to better predict and control operational costs and resource allocation by accurately forecasting token usage.

8/10

Related reading

  1. onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

    onPanda is an interactive annotation tool that lets humans correct LLM outputs token‑by‑token, then resumes generation from the corrected prefix. In a controlled study it cut median annotation time by 52% and the authors release a token‑level correction dataset (Panda‑CVL) for on‑policy fine‑tuning.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Shared Selective Persistent Memory for Agentic LLM Systems

    Apple proposes a memory architecture for agentic LLMs that selectively persists reusable context (specs, schemas, configs, constraints) across sessions and users. Shared workspaces with role‑based access and a zero‑token data‑refresh mechanism cut token usage by 97×, reduce task time by 14×, and raise task‑completion rates to 96% versus 71%‑79% for baselines.

    Apple Machine Learning Researchapple.com1 minpaper
  3. Constraint Decay: The Fragility of LLM Agents in Backend Code Generation

    A systematic evaluation of LLM agents generating multi‑file backend code shows a sharp drop in correctness when structural constraints (framework conventions, ORM usage, API contracts) are added. Across 100 tasks in 8 Python web frameworks, assertion pass rates fall ~27 points, with data‑layer bugs (bad queries, ORM violations) driving most failures. Mid‑size models cope with minimal frameworks (…

    arXiv cs.SE (Software Engineering)arxiv.org1 minpaperHN287197
  4. Optimizing Gemini 3.8 Flash for Autonomous Coding Agents: Thinking Levels and Tool Fallbacks

    The article shows how to cut latency and token waste in Gemini‑3.8‑Flash coding agents by routing each step to a suitable `thinkingLevel` (low/medium/high) based on a cheap complexity classifier, validating tool‑call payloads with Zod, and retrying failed calls with exponential back‑off. A full TypeScript harness is provided, including middleware, classifier, level mapping, and retry logic.

    SitePointsitepoint.com17 min