proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersChuxuan Hu, Hejie Cui, Norman Huang1 min readpaperadvanced

RPTune: Learned Context Curation for LLM Catalog Search

Summary

RPTune is an end‑to‑end system that learns to curate product catalogs for in‑context LLM search and fine‑tunes the LLM with a context‑relative reward. On seven real merchants it boosts search accuracy by up to 31 pp from curation and an additional ~10 pp from post‑training.

  • Encoder‑reorganizer curator orders and prunes catalog items using downstream LLM feedback.
  • Context‑relative reward fine‑tunes the LLM on curated catalogs, improving product selection.
  • Across 7 merchants and 100 complex queries each, curation adds up to 31.4 pp accuracy, post‑training adds ~10 pp on average.
  • Works with both proprietary and open‑weight LLMs, demonstrating model‑agnostic applicability.

Product engineers and IR researchers building LLM‑driven search for small‑to‑medium catalogs should care because the approach yields large accuracy gains without needing massive retrieval infrastructure.

7/10

Related reading

  1. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

    The paper presents Infinite-Parameter LLMs, where a compact hypernetwork creates feed‑forward weights from live user data and updates a Bayesian latent code online, keeping the stored model size constant while effectively having infinite parameters. This design aims to improve over standard in‑context learning and retrieval by persisting knowledge in weights and freeing context space.

    Hacker News front pagearxiv.org2 minpaperHN15743
  2. Glyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data Catalogs

    Glyph is a production system that uses coordinated LLM agents and a fine‑tuned MiniLM encoder to automatically generate column descriptions and assign ontology tags in enterprise data catalogs. It combines code‑grounded retrieval, regex, and contrastive vector search, achieving NDCG@10 0.92 and MAP@100 0.90, and provides auditable provenance for each tag.

    Apple Machine Learning Researchapple.com1 minpaper
  3. How LLMs Can Find a Needle in a Haystack

    The post explains how retrieval‑augmented generation (RAG) lets LLM‑based assistants answer questions from private corpora. It covers chunking documents into passages, embedding queries and chunks, similarity metrics, and the trade‑offs of different vector indexes (flat, IVF, HNSW). The focus is on practical design choices rather than new research.

    ByteByteGobytebytego.com12 min
  4. DataFlex-RL: An Evaluation Platform for RLVR Data Policies

    The paper introduces DataFlex‑RL, a platform to benchmark how different data‑selection policies affect reinforcement‑learning‑with‑verifiable‑rewards training. Across extensive experiments on Qwen2.5‑7B and Llama‑3.1‑8B, uniform sampling is the only method that consistently improves performance, and no alternative policy yields a statistically significant gain.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

    The paper introduces BI‑Bench, a new benchmark of real‑world BI questions derived from public dashboards, and BI‑Agent, a tool‑augmented LLM system that breaks BI workflows into search, join, and transform subtasks. Baseline LLMs hit <50 % accuracy on BI‑Bench. By orchestrating specialized data‑management tools and post‑training the model with supervised fine‑tuning and reinforcement learning on…

    Hugging Face Daily Papersarxiv.org2 minpaper