Hacker News front pageJohannes Hötter, Marko Rosenmüller, PhD13 min readintermediate
Turning GLM-5.3-Flash into a Jev-like decision model
Summary
The authors demonstrate turning an off‑the‑shelf LLM (GLM‑5.3‑Flash) into a fast, typed decision model by prompting it to output only an option index and reading the token logits. Using vLLM’s allowed_token_ids and logprob_token_ids they achieve Jev‑level accuracy and speed, even for image inputs, without any fine‑tuning.
- Number options and end the prompt with `choice_index:` so the model's first token is the option index.
- Use vLLM's `allowed_token_ids` to restrict output vocabulary and `logprob_token_ids` to get exact logits for each option.
- Token IDs for numeric options vary by tokenizer; query the model with `echo` to obtain correct IDs.
- The approach yields single‑token inference with confidence scores, matching specialized models like Jev, and works for multimodal inputs.
Backend engineers building high‑throughput LLM‑driven routing or classification services should care because this technique enables cheap, single‑token decisions with confidence scores.
6/10



