proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYazhen Xie, Xingsong Ye, Zhineng Chen1 min readpaperadvanced

Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering

Summary

The paper presents IDSpect, a reward that uses Ideographic Description Sequences to give fine‑grained feedback on Chinese character radicals for text‑to‑image models. Experiments show it improves structural quality and semantic alignment on standard benchmarks without extra inference cost.

  • IDSpect decomposes target characters into IDS tokens and aligns visual predictions at the radical level, providing deterministic, fine‑grained credit.
  • The reward operates at training time only, adding no inference‑time overhead and requiring no changes to the image generator.
  • Combined with a whole‑character semantic reward, IDSpect yields higher structural fidelity and better semantic alignment on LongText and GenTextEval benchmarks.
  • Globally unique token credit makes the reward robust to the order of detected text regions.

Engineers building text‑to‑image systems that must render complex scripts like Chinese can improve glyph accuracy without redesigning the generator.

8/10

Related reading

  1. RenderRank: Learning to Rerank Text with Compressed Visual Tokens

    RenderRank renders documents as images and uses a vision‑language model to produce compressed visual tokens for reranking, cutting input length by up to 35% while achieving higher NDCG@10 than text‑only baselines. It shows especially strong gains on long‑document datasets with half the token count and 1.7× throughput.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation

    UFO introduces an Atomized Chain‑of‑Evaluation (AEU) framework that breaks omni‑condition alignment in multi‑modal image generation into a sequential set of fine‑grained checks, achieving a 15.25 % boost in correlation with human judgments. The authors also release UFO‑Bench, a benchmark for testing how well models satisfy combined textual and visual conditions.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

    FLAT is a pre‑training framework that encodes images and text into a shared 1D token sequence with variable length via nested dropout, enabling the same embeddings for cross‑modal retrieval and generation. It attains state‑of‑the‑art scores on COCO and Flickr30K for captioning, retrieval, and text‑to‑image generation, and supports interpolation and zero‑shot composed retrieval.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. CARD: Cluster-level Adaptation with Reward-guided Decoding for Personalized Text Generation

    CARD is a hierarchical framework for personalized text generation that clusters users for group-specific LoRA adapters and learns individual preferences implicitly. It injects personalization at decoding via lightweight vectors and low-rank logit corrections, achieving superior quality and efficiency compared to baselines.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. RULER: Instance-aware Rubric Rewards for SVG Generation

    RULER introduces an instance-aware rubric reward system for generating SVG code from natural language, addressing the lack of faithful evaluation signals. It uses a VLM to score rendered SVGs against a text-derived rubric, achieving significant performance improvements over existing methods without needing ground truth or human preference data.

    Hugging Face Daily Papersarxiv.org1 minpaper