Session
Speculative Decoding Beyond Draft Models
Getting speculative-decoding speedups without training and hosting a separate draft model.
Speakers
Speculative decoding usually needs a smaller draft model to propose tokens. We survey draft-free alternatives, self-speculation, retrieval-based proposals, and n-gram guesses, and measure their real speedups against the benchmark numbers. The talk ends with guidance on which approach fits which serving budget.