A few months back, the PyTorch team released GPT, Fast (https://news.ycombinator.com/item?id=38477197), a collection of techniques to 10x the inference speed of Llama-2-7b.
This library generalizes those techniques (quantization, torch.compile, speculative decoding) to all Huggingface models, achieving an inference speed speedup of 6-7x.
As more optimization algorithms are discovered, I will post updates to the repo to HN.