A Python Library to 6-7x the inference speed of your HF models
github.com
github.com
This library generalizes those techniques (quantization, torch.compile, speculative decoding) to all Huggingface models, achieving an inference speed speedup of 6-7x.
As more optimization algorithms are discovered, I will post updates to the repo to HN.
If anyone has suggestions for other T5 inference options I would welcome them.