GPU Embedding with GGML
bloop.ai
bloop.ai
I really appreciate what GGML does but I have this irrational feeling that it'd be...unpredictable / not suited for production IMHO. (see comments here https://news.ycombinator.com/item?id=37900083 and here https://news.ycombinator.com/item?id=37898979)
I should reevaluate, this is some deep work and I assume you didn't see anything that scared you off, other than being locked to Macs for now.
See also https://github.com/rbitr/llama2.f90 which is basically the same thing but for running llama models and has 16-bit and 4-bit options and a lot more optimization.
quantized model formats:
- GGML: used with llama.cpp, outdated, support is dropped or will be soon. cpu+gpu inference
- GGUF: "new version" of the GGML file format, used with llama.cpp. cpu+gpu inference. offers 2-8bit quantization
- GPTQ: pure gpu inference, used with AutoGPTQ, exllama, exllamav2, offers only 4 bit quantization
- EXL2: pure gpu inference, used with exllamav2, offers 2-8bit quantization
here[1] is a nice overview of VRAM usage vs perplexity of different quant levels (with the example of a 70b model in exl2 format)
[1] https://old.reddit.com/r/LocalLLaMA/comments/178tzps/updated...
Thanks!
Ggml is a "framework" like pytorch etc (for the purposes of this discussion) that lets you code up the architecture of a model, load in the weights that were trained, and run inference with it. Llama.cpp is a project that I'd describe as using ggml to implement some specific AI model architectures.
examples of original models are llama(2), mistral, xwin. they are not directly related to any quantized versions. quants are mostly done by third parties (e.g. thebloke[1]).
using a full model for inference requires pretty beefy hardware. most inference on consumer hardware is done with quantized versions for that reason.
llama.cpp is a project that uses GGML the framework under the hood, same authors. Some features were even developed in llama.cpp before being ported to GGML. Ollama provides a user-friendly way to uses llama models. No ideas what it uses under the hood.
LLaMA was the model Facebook released under a non-commercial license back in February which was the first really capable openly available model. It drove a huge wave of research, and various projects were named after it (llama.cpp for example).
Llama 2 came out in July and allowed commercial usage.
But... there are increasing number of models now that aren't actually related to Llama at all. Projects like llama.cpp and Ollama can often be used to run those too.
So "Llama" no longer reliably means "related to Facebook's LLaMA architecture".
what is autoGTPTQ and exllama, what do it mean it only works with AutoGPTQ and exllama? Are those like TensorFlow Frameworks?
It looks to include submodules for GGML and GGUF from llama.cpp
That model is based on BERT and not LLaMa [2].
[1]: https://www.sbert.net/docs/pretrained_models.html
[2]: https://huggingface.co/microsoft/MiniLM-L12-H384-uncased
GPU offloading for GGUF/GGML has been available for quite a long time in Text Generation WebUI and works very well, but isn’t nearly as fast as GPTQ or the new AWQ format.
[1]: https://apple.github.io/coremltools/source/coremltools.conve...
Great work!