proomt

Search

Search posts, papers, and topics

All posts

InfoQMallika Rao26 min readtalkadvanced

Presentation: Adaptive Recommenders in the Real World: Inference, Evals, and System Design

Summary

This presentation argues that the true complexity of adaptive recommendation systems lies in their end-to-end system design, not just the machine learning models. It highlights how real-time feedback loops, retrieval freshness, multi-stage orchestration, and operational constraints enable continuous learning and evolution in production.

  • Adaptive recommenders are complex distributed feedback systems, not isolated models, with complexity at component boundaries.
  • Modern adaptive systems learn continuously and quickly, ingesting signals and evolving behavior while users interact.
  • Optimizing retrieval stages for candidate breadth and freshness often yields more impact than solely improving ranking models.
  • Effective ranking relies on rich, diverse signals including behavioral, contextual, temporal, and user intent.

Engineers building or operating large-scale recommendation or adaptive AI systems will find this valuable for understanding the critical system-level challenges beyond just model development.

7/10

Related reading

  1. Article: Beyond Relevance: A Governance-First Architecture for Enterprise Personalization

    The article proposes a governance‑first architecture for enterprise personalization, where policy‑driven steps (memory, journey graph, AI routing, scoring, trust checks, outcome simulation) shape the recommendation before it is returned. A reference FastAPI implementation demonstrates the pattern with external YAML policies and optional LLM assistance.

    InfoQinfoq.com19 min
  2. EasyPPO: Stabilizing the Critic Is Key

    EasyPPO addresses two key instability issues in PPO's critic when training large language models: biased filtering of truncated rollouts and heterogeneous return noise. It introduces actor-only overlong filtering, noise-normalized critic regression, and smaller critic mini-batches, achieving significant performance gains over vanilla PPO across various tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

    The authors show that training LLM reviewers on synthetic reviews leads to a compression of rating distributions and loss of semantic diversity, a phenomenon they call scientific-judgment collapse. They mitigate it with TrustReviewer, which uses curated training data and activation steering to preserve judgment diversity.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

    The paper presents Infinite-Parameter LLMs, where a compact hypernetwork creates feed‑forward weights from live user data and updates a Bayesian latent code online, keeping the stored model size constant while effectively having infinite parameters. This design aims to improve over standard in‑context learning and retrieval by persisting knowledge in weights and freeing context space.

    Hacker News front pagearxiv.org2 minpaperHN15743
  5. Scaling Laws for Looped Mixture of Experts

    This paper introduces Loop Scaling Laws, which jointly model recurrence and sparsity in Mixture-of-Experts (MoE) transformers. It finds that recurrence and sparsity offer complementary efficiency gains, improving prediction accuracy and enabling more efficient large model design.

    Hugging Face Daily Papersarxiv.org1 minpaperHN2