HNHacker News
TopNewBestAskShowJobs

ketchup32613

73 karma · joined May 29, 2024

submissionscomments
ketchup32613··on Do transformers need three projections? Systematic study of QKV variants
Do you want to see scaling curves wrt data and param size? I agree that 1.2B and 10B tokens is not representative, but what scale of parameters and dataset sizes would be convincing?
ketchup32613··on Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
You're not wrong about that. Speculative decoding does not affect the quality of tokens generated, as each token has to be verified by the parent model before it is output.

Each of the tokens generated by the draft model has to be verified by the parent/original model, but if this acceptance rate falls, then the speedup from speculative decoding would be eliminated. This acceptance rate, and more directly the speedup from draft models, is what "performance" refer s to in the article.