Hugging Face Daily PapersYazhen Xie, Xingsong Ye, Zhineng Chen1 min readpaperadvanced
Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering
Summary
The paper presents IDSpect, a reward that uses Ideographic Description Sequences to give fine‑grained feedback on Chinese character radicals for text‑to‑image models. Experiments show it improves structural quality and semantic alignment on standard benchmarks without extra inference cost.
- IDSpect decomposes target characters into IDS tokens and aligns visual predictions at the radical level, providing deterministic, fine‑grained credit.
- The reward operates at training time only, adding no inference‑time overhead and requiring no changes to the image generator.
- Combined with a whole‑character semantic reward, IDSpect yields higher structural fidelity and better semantic alignment on LongText and GenTextEval benchmarks.
- Globally unique token credit makes the reward robust to the order of detected text regions.
Engineers building text‑to‑image systems that must render complex scripts like Chinese can improve glyph accuracy without redesigning the generator.
8/10