D
Demerzel
Video
Speculative Decoding
Lossless inference acceleration: a small draft model proposes tokens and the target model verifies them in parallel with a rejection-sampling rule that preserves the target distribution.