proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersSiqiao Xue, Shuxuan Liu, Ning Hu1 min readpaperadvanced

ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker

Summary

ZooWork-ShopRanker is a family of open e-commerce rerankers (0.6B, 4B, 8B) trained using preference labels generated by an ensemble of LLMs. The 8B model, aligned to judge-labeled shopping preferences, significantly outperforms strong open reranker baselines on a new benchmark, ShopRank-Bench.

  • General web rerankers transfer imperfectly to e-commerce due to unique user preference signals.
  • ZooWork-ShopRanker uses an ensemble of reasoning LLMs as a preference oracle to generate training labels.
  • The 8B flagship model is distilled into more efficient 4B and 0.6B models for practical deployment.
  • ShopRank-Bench, a new 10,000-pair benchmark, measures e-commerce preference alignment.

E-commerce search engineers can leverage these open models and the new benchmark to improve product ranking by incorporating nuanced user preferences beyond topical relevance.

8/10

Related reading

  1. ML based ranking using Nrtsearch

    Yelp added an Inference Plugin to Nrtsearch that runs XGBoost and neural‑network models inside the search engine, eliminating a separate scoring service. The plugin extracts features from index documents, loads MLeap bundles from MLflow, and serves predictions on replica nodes with millisecond latency.

    Yelp Engineeringyelp.com7 min
  2. Alibaba Open Sources OpenCodeReview for AI-Assisted Code Review

    Alibaba open-sourced OpenCodeReview, an AI-powered code review CLI that combines deterministic pipelines for file selection and rule matching with an LLM agent for dynamic analysis. Used internally for two years, it claims higher precision and F1 scores than Claude Code with fewer tokens, though external reviews note recall limitations.

    InfoQinfoq.com2 min
  3. How Databricks’ marketers use data 3x more with Genie, an AI analytics assistant

    Databricks built Marge, a Genie‑powered conversational analytics assistant on a governed Marketing Lakehouse. By starting with a single high‑value use case (email campaign performance), documenting data, encoding verified answers, teaching business terminology, and embedding the tool in existing ticket workflows, they achieved 85% adoption, 3× higher data usage in decisions, 50% QoQ usage growth,…

    Databricksdatabricks.com10 min
  4. HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

    HyperBrowseComp is a new multilingual and multimodal benchmark for web-browsing AI agents, featuring 423 challenging, human-validated questions across 13 languages. It requires agents to find obscure evidence and connect information from diverse sources like videos and maps, going beyond parametric knowledge.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Modernizing the Trade Lifecycle With Governed Data and AI

    Databricks argues that modernizing the trade lifecycle now hinges on building a governed, real‑time data foundation that spans research, trading, risk, ops and compliance, rather than isolated AI pilots. Starting with a few high‑value questions—execution cost, shock risk, exception rates—and using Unity Catalog and Agent Bricks lets firms achieve measurable speed and auditability gains before sca…

    Databricksdatabricks.com5 min