Hacker News front pageByteShape17 min readintermediate
Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
Summary
ByteShape releases full ShapeLearn quantizations for Qwen 3.8 27B, showing that their GPU‑specific GGUFs (GPU‑1…GPU‑5) dominate the quality‑throughput frontier across six GPUs, with GPU‑5 hitting 99.63 % of BF16 accuracy at 90 TPS on a 13.1 GB model. Speculative decoding (MTP, DFlash2) further boosts throughput, and the Lite set remains competitive.
- ShapeLearn‑Lite (quick‑turn quantization) holds up well; the full ShapeLearn run improves both accuracy and speed for all five model sizes.
- GPU‑5 (IQ4_XS, 3.84 bpw, 13.1 GB) is the default recommendation, achieving 99.63 % of BF16 aggregate score; GPU‑4 offers a smaller 11 GB variant with 98.72 % of BF16.
- Speculative decoding with MTP (multimodal‑capable) or DFlash2 (text‑only, faster) consistently raises tokens‑per‑second on every GPU tested.
- Across RTX Pro 6000 (96 GB) and RTX 5090 (32 GB) GPUs, ShapeLearn models sit on the measured quality‑throughput frontier, outperforming competing quantizations from Unsloth, ISTA‑DASLab, AtomicChat, and others.
Quantized LLMs are the only way to run 27 B‑parameter models on consumer‑grade GPUs. ByteShape’s ShapeLearn quantizations demonstrate that careful, GPU‑aware PTQ can approach full‑precision quality while fitting into 11‑13 GB VRAM, making high‑quality inference accessible without expensive hardware…
6/10
.png)


