https://github.com/microsoft/azurechatgpt
Past discussion:
https://github.com/microsoft/azurechatgpt
Past discussion:
There are some great open-source projects in this space – not quite the same – many are focused on local LLMs like Llama2 or Code Llama which was released last week:
- https://github.com/jmorganca/ollama (download & run LLMs locally - I'm a maintainer)
- https://github.com/simonw/llm (access LLMs from the cli - cloud and local)
- https://github.com/oobabooga/text-generation-webui (a web ui w/ different backends)
- https://github.com/ggerganov/llama.cpp (fast local LLM runner)
- https://github.com/go-skynet/LocalAI (has an openai-compatible api)
- https://github.com/trypromptly/LLMStack (build and run apps locally with LocalAI support - I'm a maintainer)
The UI is relatively mature, as it predates llama. It includes upstream llama.cpp PRs, integrated AI horde support, lots of sampling tuning knobs, easy gpu/cpu offloading, and its basically dependency free.
GPTQ has also been merged into Transformers library recently ( https://huggingface.co/blog/gptq-integration ).
GGML quantization format used by llama.cpp also supports (8,6,5,4,3, and 2 bit quantization).
On a related note it doesn't seem like many local runners are leveraging techniques like PagedAttention yet (see https://vllm.ai/) which is inspired by operating system memory paging to reduce memory requirements for LLMs.
It's not quite what you mentioned, but it might have a similar effect! Would love to know if you've seen other methods that might help reduce memory requirements.. it's one of the largest resource bottlenecks to running LLMs right now!
The hint for me is that the models compress so well, that suggests the information content is much lower than the size of the uncompressed model indicates which is a good reason to investigate which parts of the model are so compressible and why. I haven't looked at the raw data of these models but maybe I'll give it a shot. Sometimes you can learn a lot about the structure (built in or emergent) of data just by staring at the dumps.
Full Disclosure: This is my tool
Normally we ban accounts that do nothing but promote their own links, but as you've been an HN member for years, I'm not going to ban you, but please do stop doing this! We want people to use HN to read and post things that they personally find intellectually interesting—not just to promote something.
If I go back far enough (a couple hundred comments are so), it's clear that you used to use HN in the intended spirit, so this should be fairly easy to fix.