Related reading
Generate fonts where every LLM token is the same width
This article introduces "token-space fonts" which render each LLM token with equal width, allowing users to visualize how text is tokenized. It provides a web-based compiler to generate these fonts for various LLM tokenizers and instructions for integrating them into Discord and Slack.
Hacker News front pagemesh.host7 minHN9423Labeled matches: why is this not in every regex engine?
The author shows how to label tokens (dates, money, emails, etc.) with a single regex pass using extended operators (`&` for intersection, `~` for complement) in the resharp library. A tiny benchmark compares 10 patterns against spaCy’s NER, reporting ~1.9 GB/s (≈4500× faster) on 8 threads. The post lists the concrete patterns and argues that for deterministic, regular‑language entities regex can…
Beyond Selection: Token Parameterization for Extreme Visual Token Compression
This paper introduces Braco, a lightweight coder for extreme visual token compression in vision-language models. It uses a novel token parameterization approach to maintain visual grounding and achieve high accuracy at 23x-64x compression, significantly improving efficiency over prior methods.
Hugging Face Daily Papersarxiv.org1 minpaperSrijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts
Srijika is a system for restyling OpenType fonts across nine Indic scripts by reusing existing template font layout data (GSUB, GPOS). This approach ensures complex script features like conjuncts remain consistent, addressing a major challenge in Indic font generation.
Hugging Face Daily Papersarxiv.org1 minpaper

