ubermenchh··on Show HN: Mini-vLLM in ~500 lines of Pythonyes it does continous batching along with paged attention and prefix caching. i am also goint to be adding some more inference techniques
ubermenchh··on Show HN: Mini-vLLM in ~500 lines of PythonHaha, i just wanted my repo to be out here. If someone finds it interesting they can always just check the repo. And you're close, its about getting faster responses from the model by manipulating the request queues and memory.