proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYuta Oshima, Ku Onoda, Yusuke Iwasawa2 min readpaperadvanced

AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation

Summary

AutoRef uses a coding LLM to automatically rewrite the harness that orchestrates multi‑reference image generation, keeping the generator frozen. The resulting harness lifts FLUX.2‑4B scores from 5.72 to 7.37 on MultiBanana and generalizes across models and benchmarks.

  • A coding agent iteratively rewrites harness code, separating feedback (proposal) and selection tasks, and continues search from a beam of top‑ranked harnesses.
  • Optimizing only the harness (no model fine‑tuning) improves FLUX.2‑4B performance on four‑reference MultiBanana from 5.72 to 7.37, matching proprietary systems.
  • The discovered AutoRef‑Harness transfers to different generators, numbers of references, evaluators, and reasoning models without re‑optimization.
  • The method treats harness design as a searchable program space, enabling systematic improvement over hand‑crafted harnesses that vary widely in quality.

Teams building multi‑reference generation pipelines can boost image quality without costly model retraining by automating harness design.

8/10

Related reading

  1. FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

    FuseReg replaces the fixed heuristic of selecting encoder layers for representation autoencoders with a regularization that trains on random subsets of layers, making the downstream decoder robust to any fusion. This yields higher reconstruction quality (PSNR) and lowers unguided generation FID by up to 29% on ImageNet‑256, all without changing the pretrained visual encoder.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Recursive self-improvement of AI research agents

    The paper introduces AIDE², an AI research agent that rewrites its own code, benchmarks each version, and adopts the best performing changes—a process they call recursive self‑improvement. In an 8‑day autonomous run it produced seven improvements that beat a strong human‑engineered baseline on four unseen benchmarks and reduced reward‑hacking from 55 % to 32 %.

    Hugging Face Daily Papersarxiv.org2 minpaperHN31
  3. Automating coherent long-form video generation

    Google Research introduces a unified multi-agent framework to autonomously generate temporally consistent, long-form video narratives. It overcomes identity drift and cascading failures of current linear AI pipelines by treating generation as a global optimization and world-state tracking problem.

    Google Researchresearch.google10 min
  4. Reasoning with Image Generation

    ReImaGin proposes using image generation models as a flexible visual reasoning mechanism for multimodal LLMs, moving beyond rigid visual tools. It consistently outperforms text-only and specialist vision-tool baselines by up to 25% on diverse visual reasoning tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper