proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersGuanzheng Chen, Viet Dac Lai, Subhojyoti Mukherjee1 min readpaperadvanced

DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation

Summary

DISCO is a distributed architecture that disaggregates grounding from reasoning in LLMs to combat "context rot" in long contexts. It uses Worker LLMs for parallel grounding and a Driver LLM for orchestration, maintaining high accuracy and reducing inference costs by over 80% on million-token inputs.

  • Context rot describes the collapse of LLM reasoning quality as input context windows grow very long.
  • DISCO disaggregates the search-heavy contextual grounding from complex reasoning tasks.
  • It employs a distributed setup with Worker LLMs for parallel, localized grounding and a central Driver LLM for orchestration and reasoning.
  • The Driver LLM is trained via Reinforcement Learning (GRPO) to optimize query planning and evidence reduction.

Engineers building applications with long-context LLMs should care about DISCO as it offers a robust, cost-effective solution to the context rot problem, improving reliability and performance.

8/10

Related reading

  1. Dynamically Scaled Activation Steering

    Dynamically Scaled Activation Steering (DSAS) is a method‑agnostic framework that learns per‑token, per‑layer scaling factors to turn existing activation‑steering interventions on only when a model is likely to produce undesired output (e.g., toxic text). The scaling can be optimized jointly with any steering function, improves the toxicity‑utility trade‑off on language models, transfers to text‑…

    Apple Machine Learning Researchapple.com1 minpaper
  2. Verifiable Social Reasoning for LLM Assistants

    The paper introduces Fuse, a multi‑agent simulation that gives LLM assistants a verifiable ground‑truth task for social reasoning by hiding a target agent’s motive and letting a user‑mediated conversation infer it. Experiments on 12 LLMs show user mediation makes reasoning harder, models are biased by user framing, need more detail than humans, and longer chats don’t always help.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Shared Selective Persistent Memory for Agentic LLM Systems

    Apple proposes a memory architecture for agentic LLMs that selectively persists reusable context (specs, schemas, configs, constraints) across sessions and users. Shared workspaces with role‑based access and a zero‑token data‑refresh mechanism cut token usage by 97×, reduce task time by 14×, and raise task‑completion rates to 96% versus 71%‑79% for baselines.

    Apple Machine Learning Researchapple.com1 minpaper
  4. How Data 360 Builds Trusted Context: The Enduring Layer for Enterprise AI

    Salesforce’s Data 360 provides a shared runtime that assembles the minimal, authorized slice of enterprise data (“Trusted Context”) for each AI‑agent turn. A six‑stage Agent Context Engine (Resolve, Plan, Reconcile, Govern, Compile, Learn) pulls data from structured, unstructured, and streaming sources across Salesforce, Snowflake, Databricks, etc., applies fine‑grained policy, and returns a toke…

    Salesforce Engineeringsalesforce.com11 min
  5. Constraint Decay: The Fragility of LLM Agents in Backend Code Generation

    A systematic evaluation of LLM agents generating multi‑file backend code shows a sharp drop in correctness when structural constraints (framework conventions, ORM usage, API contracts) are added. Across 100 tasks in 8 Python web frameworks, assertion pass rates fall ~27 points, with data‑layer bugs (bad queries, ORM violations) driving most failures. Mid‑size models cope with minimal frameworks (…

    arXiv cs.SE (Software Engineering)arxiv.org1 minpaperHN287197