1,780 karma · joined January 31, 2014
On a related note it doesn't seem like many local runners are leveraging techniques like PagedAttention yet (see https://vllm.ai/) which is inspired by operating system memory paging to reduce memory requirements for LLMs.
It's not quite what you mentioned, but it might have a similar effect! Would love to know if you've seen other methods that might help reduce memory requirements.. it's one of the largest resource bottlenecks to running LLMs right now!
There are some great open-source projects in this space – not quite the same – many are focused on local LLMs like Llama2 or Code Llama which was released last week:
- https://github.com/jmorganca/ollama (download & run LLMs locally - I'm a maintainer)
- https://github.com/simonw/llm (access LLMs from the cli - cloud and local)
- https://github.com/oobabooga/text-generation-webui (a web ui w/ different backends)
- https://github.com/ggerganov/llama.cpp (fast local LLM runner)
- https://github.com/go-skynet/LocalAI (has an openai-compatible api)
Continue also works with various backends and fine-tuned versions of Code Llama. E.g. for a local experience with GPU acceleration on macOS, continue can be used with Ollama (https://github.com/jmorganca/ollama):
ollama pull codellama
from continuedev.src.continuedev.libs.llm.ollama import Ollama
config = ContinueConfig(
models=Models(
default=Ollama(model="wizardcoder:34b-python")
)
)The download (both as a Mac app and standalone binary) is available here: https://github.com/jmorganca/ollama/releases/tag/v0.0.16. And I will work on getting that brew formula updated as well! Sorry to see you hit an error!
ollama run phind-codellama "write c code to reverse a linked list"
To run this on an m1 Mac or similar machine, you'll need around 32GB of memory for the 4-bit quantized version since it's a 34B parameter model and is quite big (20GB).There's a "PrivateGPT" example in there that is similar to your third point above: https://github.com/jmorganca/ollama/tree/main/examples/priva...
Would love to know your thoughts
curl -X POST http://localhost:11434/api/generate -d '{
"model": "codellama",
"prompt":"write a python script to add two numbers"
}' ollama pull codellama:7b-instructThis should be fixed now! To update you'll have to run:
ollama pull codellama:7b-instruct ollama run codellama "write a python function to add two numbers"
More models coming soon (completion, python and more parameter counts)It's especially interesting how you could combine different model types - e.g. translation + text completion (or image generation) – it could be a pretty powerful combination...
A few folks and I have been building a tool with it in Go for pulling & running multiple models, and serving them on a REST API: https://github.com/jmorganca/ollama
In similar light, you haven't checked it out, llama.cpp also has a pretty extensive "server" tool (in its examples directory in the repo) with a web ui and support for grammar (e.g. forcing the output to be JSON)
A few more:
- Wizard Vicuna 13B uncensored
- Nous Hermes Llama 2
- WizardLM Uncensored llama2
Over the weekend Lance Martin got it working with local models using llama.cpp and ollama.ai which saves $ on longer sims since all inference happens locally https://twitter.com/RLanceMartin/status/1690829179615657985. It's neat how the AI agents interface with each other – e.g. one will host a party and invites will be sent throughout the group
IIRC chat bots are central the vision Facebook has with LLMs (e.g. every instagram account has a personal chat bot), so I would expect the Llama models to get increasingly better at this task.
That said the 7B and 13B models definitely don't quite seem ready yet for production customer interaction :-)
If you're looking to try the "open" models like Llama 2 (or it's uncensored version Llama 2 Uncensored), check out https://github.com/jmorganca/ollama or some of the lower level runners like llama.cpp (which powers the aforementioned project I'm working on) or Candle, the new project by hugging face.
What's are folks' take on this vs Llama 2, which was recently released by Facebook Research? While I haven't tested it extensively, 70B model is supposed to rival Chat GPT 3.5 in most areas, and there are now some new fine-tuned versions that excel at specific tasks like coding (the 'codeup' model) or the new Wizard Math (https://github.com/nlpxucan/WizardLM) which claims to outperform ChatGPT 3.5 on grade school math problems.
curl https://ollama.ai/v2/_catalog | jq
Then to list "tags" for a given model (e.g. llama2): curl https://ollama.ai/v2/library/llama2/tags/list | jqA maintainer of the project has been collecting a full list here (with different quantization levels), most of which are Llama 2-based: https://gist.github.com/mchiang0610/b959e3c189ec1e948e4f6a1f...
Since the release of Llama 2 the number of models based on it has been growing significantly.. some popular ones:
- codeup (A code generation model - DeepSE)
- llama2-uncensored (George Sung)
- nous-hermes-llama2 (Nous Research)
- wizardlm-uncensored (WizardLM)
- stablebeluga (Stability AI)
The article also recommends oobabooga's text-generation-webui which includes a full web dashboard.
There's also the fact that the data being sent to these LLM "destinations" could be significantly more valuable (or contain significantly more sensitive information) than the average segment identity or track objects.
https://python.langchain.com/docs/integrations/llms/ollama
This can be a great option if you'd like to keep your data local versus submitting it to a cloud LLM, with the added benefit of saving costs if you're submitting many questions in a row (e.g. in batches)