I do have a feeling of dejavú, like I've seen this before on hn.
I do have a feeling of dejavú, like I've seen this before on hn.
You're either thinking of speculative encoding more generally, or Medusa: https://arxiv.org/abs/2401.10774
Therefore, unsurprisingly, the cost per item of inference on a batch of items is significantly lower when the batch is e.g. 8 than 1 (in the case of Transformers there are further gains to be made because roughly half of the attention calculations in token k+1 are identical to the calculations of token k and can be easily reused by writing the formulas a certain way, the keyword to look for is causal attention mask).
In any reasonable GPU inference setup the weights would be preloaded.
The main reason I can come up with why doing the same calculation 8 times in parallel instead of 8 times sequentially is that you benefit from better locality of reference.
From the below Modal link they use: >continuous batching, so multiple generations can take place at the same time on a single container
>PagedAttention, which applies memory paging to the attention mechanism’s key-value cache, increasing throughput
[1]https://modal.com/docs/examples/text_generation_inference