GGUF, the Long Way Around
vickiboykis.com
vickiboykis.com
While I haven't seen the model storage and distribution format, the rewrite to GGUF for file storage seems to have been a big boon/boost to the project. Thanks Phil! Cool stuff. Also, he's a really nice guy to boot. Please say hi from Fern to him if you ever run into him. I mean it literally, make his life a hellish barrage of nonstop greetings from Fern.
It is hard to keep metadata minimal, and before long, you will start to have many different "atom"s and end-up with things that mov supports but mp4 doesn't etc etc. (mov format is generally well-defined and easy-to-parse, but being a binary format, you have to write your parser etc is not a pleasant experience).
If you just want minimal dependency, flatbuffers, capnproto, json are all well-supported on many platforms.
[1] https://github.com/ggerganov/llama.cpp/blob/master/ggml-cuda...
GG is Georgi Gerganov
https://www.bleepingcomputer.com/news/security/malicious-ai-...
it enables really clean separation of the core autodiff library and whatever backend you want to use to accelerate the graph computations, which can simply read the file and be completely independent of the core implementation
but also, if you just store the tensors in some arbitrary order and then store the indices of the order in which they have to read and traversed, you can easily adjust the graph to add stuff like layer fusion or smth similar (i'm not really familiar w/ comp graph optimisations tbh)
what would an alternative look like anyway?
llama
mpt
gptneox
gptj
gpt2
bloom
falcon
rwkvI'm also a bit confused by the quantization aspect. This is a pretty complex topic. GGML seems to use 16bit as per the article. If was pushing it to 8bit, I reckin I'd see no size improvement the GGML file? The article says they encode quantization versions in that file. Where are they defined?
In designing software, there's often a trade off between (i) generality / configurability, and (ii) performance.
llama.cpp is built for inference, not for training or model architecture research. It seems reasonable to optimize for performance, which is what ~100% of llama.cpp users care about.