Conceptual

Speculative Decoding

Lossless inference acceleration: a small draft model proposes tokens and the target model verifies them in parallel with a rejection-sampling rule that preserves the target distribution.