
YL
Yanxiao Liu, Sicheng Wan, Zhan Gao, Deniz G\"und\"uz
· 1 min read
ResearcharXiv cs.LG
Watermarkable Multi-Draft Speculative Sampling via Poisson Processes
arXiv:2609.21858v1 Announce Type: cross
Abstract: Large language models (LLMs) have achieved state-of-the-art performance across a wide range of tasks, motivating two important aspects of deployment: inference efficiency and output provenance, which can be tackled by speculative sampling and watermarking, respectively. However, recent works have shown that combining these two goals is highly nontrivial and can be potentially impossible. In this work, we develop a novel multi-draft speculative sampling algorithm based on Poisson processes that improves the frontier of this fundamental trade-off. The proposed algorithm has strong sampling efficiency on its own and, more interestingly, is naturally watermarkable: we can embed an unbiased watermark without degrading speculative acceptance. Moreover, our algorithm is based on an exact list-coupling-without-communication scheme, which yields a drafter invariance property that benefits both sampling and watermarking. It is the first multi-draft, drafter-invariant speculative sampling scheme that maintains both watermark strength and sampling efficiency, and we experimentally verify its strong performance in both aspects.
Original source
This story was published by arXiv cs.LG and written by Yanxiao Liu, Sicheng Wan, Zhan Gao, Deniz G\"und\"uz. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


