HNHacker News
TopNewBestAskShowJobs

MDK8888

42 karma · joined January 5, 2024

submissionscomments
MDK8888··on A Minimal Implementation of Vllm
I built vllm from the ground up. Starting from the same kernels, I integrate them into GPT2, build the KV cache manager, the scheduler, and finally the FastAPI server on top. This is purely educational for now, but in the future it will become production ready.
MDK8888··on A Python Library to 6-7x the inference speed of your HF models
This has been added to the Readme under the 'Background' section.
MDK8888··on A Python Library to 6-7x the inference speed of your HF models
Hey, hope you're doing well! I've updated the repo with instructions on how to get started and documentation!
MDK8888··on A Python Library to 6-7x the inference speed of your HF models
A few months back, the PyTorch team released GPT, Fast (https://news.ycombinator.com/item?id=38477197), a collection of techniques to 10x the inference speed of Llama-2-7b.

This library generalizes those techniques (quantization, torch.compile, speculative decoding) to all Huggingface models, achieving an inference speed speedup of 6-7x.

As more optimization algorithms are discovered, I will post updates to the repo to HN.

MDK8888··on Show HN: SageMode - A Python library for deploying, scaling, and monitoring LLMs
SageMode is an open source python library for deploying, scaling, and monitoring your LLMs at scale. It's designed to be highly flexible but also standardized. You can deploy any Huggingface Model that exists in version 4.26 or earlier as well as any PyTorch Model. If you find any issues with it, open an issue-I'd love to iron out any problems.