Load LLaMA Models Instantly
twitter.com
twitter.com
For me personally this matters most because right now when llama.cpp runs out of the 2048 tokens it segfaults and this causes difficulties. In interactive mode if it goes off rails and generates 1000 tokens of nonsense then that nonsense is taking up tokens for the next line from chat. In normal mode where it just runs once and all history has to be manually supplied this can be avoided.
jart is on another plane of existence—and 100% earned this flex
Generated code practically always runs correctly the first time.
I'm not a particularly proficient and before this I would probably be happy with averaging 30 LOC/hr.
[1] https://github.com/ggerganov/llama.cpp/commit/5b8023d9354010...
I ran Alpaca 7B Q4 almost instantly because they provided Curl's to download it. Super simple. But it seems most aren't doing that because it's prone to getting Facebook's gaze. So.. what's recommended?
I happened to find this[2], but i think that's the non-quantized raw models? Not sure yet.
[1]: Won't bother with 65B, can't fit in memory i believe? [2]: https://github.com/shawwn/llama-dl/blob/main/llama.sh
edit: I forgot about https://github.com/cocktailpeanut/dalai - i suspect this is best in breed atm? Though a Docker container would be nice to wrangle all the dependencies
I'm using a not-that-new Macbook Pro (intel, 16GB memory) and was able to run the 7B and 13B models that way, tried 30B but it seemed to hang.
A bit heavy handed perhaps to use NPM/etc, but the Dockerfile really helps me ignore all the dependencies i'm adding hah.
In this case, it looks like jart@ modified malloc to capture the memory generated by the loading process and serialized that to disk. When you run, the application calls mmap to make a virtual memory association with the bytes on disk- so any time you access RAM and it's not yet loaded, ti gets loaded from disk. At that point it gets saved by the kernel in a page cache and since the files on disk don't change, those pages can stay in memory longer than the process. So when the process restarts, all those RAM requests are immediately mapped to already-cached virtual memory, rather than reading from disk.
The inference library here supports a data pointer that would point to the memory mapped location.
This is faster than relying on the kernel's disk read cache; in that case, you'd still need to convert the data from the disk format to the in-memory format.
Normally the data build process is run as an external program that writes the mmap-ready structure to disk (an example is the BLAST program which writes the DNA sequence data into an index structure that is mmapped at runtime). But in this case ti looks like using an instrumented malloc() helps simplify the process of building the disk structure.
That's because the people that develop these models are often data scientists and have little to no experience with systems programming, optimizations, etc.
It’s really simple to take some python inference code and wrap FastAPI around it.
However, inference servers exist for a reason. You’ll quickly find that performance, VRAM usage, model management, etc isn’t practical with the FastAPI approach.
Speaking personally inference server implementations like Nvidia Triton bring a model to performance metrics that are absolutely night and day vs the FastAPI approach - in many cases orders of magnitude higher performance in terms of response time and requests per second.
- Dynamic batching while limiting latency to a set threshold
- Running multiple instances of a model, effectively load-balancing inference requests.
- Loading/unloading/running multiple versions of models dynamically, which is useful if you want to update (or roll back) your model while not interfering with existing inference requests.
Its client provides async based inference APIs, so you can easily put a FastAPI-based API server in front and don't necessarily need a queue (like Celery).
FastAPI loads a model statically on startup. There are some hacks to reload versions and new models via things with load balancers, etc but they’re just that - hacks. There are also known issues with TensorFlow especially having poor memory management over request count.
FastAPI is great but at the end of the day it’s Python and the performance reflects that (more on this later).
With Nvidia Triton you get:
- Automatic support for various model frameworks/formats: native PyTorch/TensorFlow, ONNX, and more.
- Dynamic batching. You can configure an SLA with max additional latency for response time where Triton will queue requests from multiple clients over a given time period and pass them through though model batched. If you have the VRAM (you should) it’s an instant performance multiplier.
- Even better performance: Triton can do things like automatically compile/convert a model to TensorRT on the runtime hardware. This allows you to deploy models across hardware families with optimized performance while not worrying about the specific compute architecture or dealing with TensorRT itself.
- Optimized and efficient use of multiple GPUs.
- Model version management. Triton has a model management API you can use to upload a new model/version and load it dynamically. It can hot load/reload a model and serve it instantly, with configuration options for always serving the latest model or allowing client to request a specific version.
- Performance metrics. It has built in support for Prometheus.
- Other tools like Model Navigator and Performance Analyzer. You can pass a model to these tools and they will try every possible model format, batch size, etc, etc against an actual Triton server and produce a report and optimized model configuration based on your selected parameters - requests per second, response time, etc. Even memory/compute utilization, power usage, and more.
- Out of the box without any of these tricks Triton is faster, uses less memory, less GPU compute, and less CPU compute. Written in C and optimized by Nvidia.
- It’s a single implementation (often container) that from the get go is smaller, lighter weight, and easier to manage than pip installing a bunch of dependencies and the entire runtime framework itself. It exists solely to serve models and serve them well.
When you add it up (as I mentioned) I’ve personally seen cases where requests per second increase by orders of magnitude with lower response times than a single request against FastAPI (or similar). Plus all of the mlops and metrics features.
Frankly, it’s pretty amazing.
The reason is that the kernel gives you this featuer and it's really powerful, so why not take advantage of it?
During dev work you often want an easily restarted stack (code changes). Anyway, if you use this approach, you can just have the python server stay resident and then have another app mmap it (shared memory) instead of doing inference over an API, whcih is always awkward.
> the gains here are mostly due to not copying memory anymore, and better cooperation with the kernel's page manager. We unfortunately aren't getting any additional gains from lazy page loading, since this is a dense model. To generate a single token, every single page in the model file needs to be loaded. What this means is that first runs that load from spinning disk are still going to be slow, even though the average case has greatly improved.
If I understand correctly, this only provides a speed up when the model is already in the OS’s file system cache, and in any case you still have to load the entire model into memory.
But if the intent is to have multiple processes share the samedata then mmap likely is the better choice.
With a GDDR6X memory bandwidth of 936 GB/s and PCIe 4.0 x16 bandwidth of 64GB/s loading something like 20GB into the VRAM of an RTX 3090 shouldn't take longer than 1/2 a second or so, right? (assuming it is in the kernel cache)
https://github.com/ggerganov/llama.cpp/issues/91#issuecommen...
main() {
magic_init(); // replaces malloc() with ./magic.dat file
if (magic->was_committed) {
model = magic->model; // use the heap built from last run
} else {
// create a new heap
model = new model;
llama_model_load(model); // pre-existing model loading code
magic->model = model;
magic_commit(); // mark memory transaction complete
}
// rest of pre-existing code
}
You could write a Python extension for this code https://github.com/ggerganov/llama.cpp/commit/5b8023d9354010... and see if it works for your use case.Or am I missing something?
The issue is mainly stack-allocated memory. For example, our `model` object here needed to be allocated on the heap, rather than being declared as an automatic variable. Since automatic objects is really an optimization unique to low-level languages like C++, I doubt that's going to be an issue for something like Python, where everything is probably a heap object.
Oh! Also be careful about .data objects (i.e. static variables). Those could prove problematic too. Although my code could easily be updated to memcpy() the memory between _etext and _end into the file.
Thanks!