What is speculative decoding?
Maor Ashkenazi, research team lead at NVIDIA, explains how speculative decoding speeds up language model inference by drafting tokens ahead and having the full model verify or correct them.
Maor Ashkenazi, research team lead at NVIDIA, explains how speculative decoding speeds up language model inference by drafting tokens ahead and having the full model verify or correct them.