1
How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus
This paper investigates the "lossless" claim of Orthrus, a hybrid architecture for accelerating LLM inference. It finds that under BF16 precision, Orthrus diverges from the exact autoregressive output trajectory in over 50% of cases, though FP32 maintains exact matching. Despite BF16 divergence, downstream task performance was not systematically degraded.
Hugging Face Daily Papersarxiv.org1 minpaper
