Hugging Face Daily PapersChaoqian Ouyang, Ling Yue, Libin Zheng1 min readpaperadvanced
TokenCast: Forecasting Token Consumption During LLM Agent Execution
Summary
LLM agent token consumption varies widely and is hard to predict due to dynamic execution and growing context. TokenCast proposes a composable cost representation for execution segments to forecast token usage, reducing prediction error and enabling better budget control.
- LLM agent token consumption can vary by over an order of magnitude for the same task.
- TokenCast learns a composable cost representation for each execution segment, including its own consumption and context growth.
- It accounts for cumulative context growth, which inflates input size for subsequent LLM calls.
- TokenCast reduces mean absolute error by 14.5% against strong comparators across various tasks and models.
Engineers deploying LLM agents can use this to better predict and control operational costs and resource allocation by accurately forecasting token usage.
8/10


