1
Online Learning with LLM Experts from Limited Feedback
The paper models prompt routing to multiple LLM experts as a bandit problem with limited feedback and proposes algorithms that achieve sublinear regret in both full‑information and bandit settings. Experiments demonstrate that the methods learn effective routing strategies across diverse LLMs using only a small feedback budget.
Hugging Face Daily Papersarxiv.org2 minpaper
