BASS: Batched Attention-optimized Speculative Sampling

Haifeng Qian,Sujan Kumar Gonugondla, Sungsoo Ha,Mingyue Shang,Sanjay Krishna Gouda, Ramesh Nallapati,Sudipta Sengupta,Xiaofei Ma,Anoop Deoras

arxiv（2024）

引用 0|浏览0

暂无评分

摘要

Speculative decoding has emerged as a powerful method to improve latency and throughput in hosting large language models. However, most existing implementations focus on generating a single sequence. Real-world generative AI applications often require multiple responses and how to perform speculative decoding in a batched setting while preserving its latency benefits poses non-trivial challenges. This paper describes a system of batched speculative decoding that sets a new state of the art in multi-sequence generation latency and that demonstrates superior GPU utilization as well as quality of generations within a time budget. For example, for a 7.8B-size model on a single A100 GPU and with a batch size of 8, each sequence is generated at an average speed of 5.8ms per token, the overall throughput being 1.1K tokens per second. These results represent state-of-the-art latency and a 2.15X speed-up over optimized regular decoding. Within a time budget that regular decoding does not finish, our system is able to generate sequences with HumanEval Pass@First of 43 single-sequence speculative decoding. Our peak GPU utilization during decoding reaches as high as 15.8 and around 10X of single-sequence speculative decoding.

查看译文

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要