J
jeremy
Video
Speculative Decoding in High Batch LLM Inference using KV Cache Amortization on H100 GPUs
Speculative decoding operates within high-performance computing and large language model inference by utilizing a tandem mechanism where a low-cost draft model generates multiple candidate tokens in …