Ask HN: What's the best self hosted/local alternative to GPT-4?
[^1]: https://news.ycombinator.com/item?id=36134249
[^1]: https://news.ycombinator.com/item?id=36134249
You’re stuck with openai, and you’re stuck with whatever rules, limitations or changes they give you.
There are other models, but specifically if you’re actively using gpt-4 and find gpt-3.5 to be below the quality you require…
Too bad. You’re out of luck.
Wait for better open source models or wait patiently for someone to release a meaningful competitor, or wait for openai to release a better version.
That’s it. Right now, there’s no one else letting people have access to their models which are equivalent to gpt-4.
The design of abstractions, prompt engineering, custom fine-tunes and software engineering required to ship a valuable application on top of that interface counts as "building an app" in my book.
I don’t think OpenAI’s contribution should be so understated- they built a technology that was considered science fiction just a few years ago. They deserve all the credit for the “AI”.
I would also add "distribution" to the "app and UX" part, but you can certainly build a valuable business upon an API "everyone else has access to" - plenty of companies out there that do that.
[0]https://twitter.com/englishpaulm/status/1623701758781558784
We can now write software that interprets language under the hood (to some degree). The value propositions enabled by this change in the world are so vast, and partly so complex - to make absolute statements like "yeah but you don't control the model, so anyone can copy your solution" seems out of touch to me. What subset of technology doesn't get reverse engineered? Either this applies almost nowhere (because every piece of tech that an engineer can get their hands on is effectively open), or everywhere.
That doesn't mean that everything else in the stack is window dressing though - custom, domain specific wrangling with the different api endpoints, finding a satisfying prompt, temperature param etc. for specific tasks - the entire process of designing systems around an LLM-api has many intricacies to it, lots of which are completely uncharted territory.
I can assure you: very smart people are knees deep in this process, and they deserve the credit for their share of the value that is being created.
However, the "very very temporary space" might as well lead to momentum and a moat in a subdomain, and anyone who met a sufficiently large number of smart people knows that lots of them are very pragmatic, don't chase prestige, and enjoy laying ground work for future iterations.
It's certainly building an app. It's not building an AI app, though. It's building a front-end to an existing AI application.
https://cloud.google.com/vertex-ai/docs/generative-ai/models...
You can generate the training data for this with 3.5 and 4 and tune smaller models with the resulting data. For lots of tasks, this results in robust results, which btw are also faster than 3.5 turbo.
Replace social media / graph APIs with the ones from OpenAI
So I have to believe that the people making these products are intending to make as much cash as possible up front and aren't aiming for a long-term thing.
Depending on your business case, the 4096 tokens given to you have to go quite far. Vector embeddings are not "easy" to work with. Trying to splat together a range of techniques to craft a good prompt is hard™.
Adding in Actions (e.g. using headless browsers to open pages etc) is also pioneering territory.
Sucks that OpenAI currently has the market, but there is still plenty of reasons to develop on top of it.
The technology is just GPT and transformers, the open source alternatives are just not as advanced as the GPT4 model yet. It might change with current trajectory, mainly because OpenAI keeps nerfing it.
Claude should be more well known, imo. ChatGPT/GPT4 gets all the hype, but Claude is really good too... sometimes even better.
And this is Claude+: https://poe.com/Claude%2B
So for example Tesla has/had a lead in EV tech, but their supercharger network is a moat because it’s very difficult for a rival to compete with, even if they have equivalent EV tech.
I know there isn’t established terminology for this. It is a nitpick, but I think a lead is already a term we have, and moat has connotations of being in some way ‘unfair’. To me, just creating a better product isn’t unfair in any meaningful way, while moats in some circumstances can represent anti trust issues.
But perhaps at this stage Google's not meant to / need to build cutting edge products, but rather focus on commoditizing the ones that prove profitable.
How is Google "losing" when they are not even in the same market? Are you expecting Google to develop and sell access to an LLM API or use AI internally to enhance it's other products (which it jas been doing to great effect)
OpenAI has a lead on Bard, LLaMA and friends, but I would expect that to close in the next few months or years.
I also generally prefer Claude or Claude+ over GPT4 or ChatGPT (GPT 3.5) for explanations too, which tend to be more thorough without any special prompting.
There’s a gap, it’s closing, likely faster than anticipated.
Huggingface awaits :)
A quick test of the huggingface demo gives reasonable results[1]. The actual model behind the space is here[2], and should be self-hostable with reasonable effort.
0. https://arxiv.org/abs/2305.14314 1. https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi 2. https://huggingface.co/timdettmers/guanaco-33b-merged
CPU Version: https://huggingface.co/TheBloke/guanaco-65B-GGML
GPU Version: https://huggingface.co/TheBloke/guanaco-65B-HF
4bit GPU Version: https://huggingface.co/TheBloke/guanaco-65B-GPTQ
How much ram and vram does one need to run 4,13,33,65B models at a reasonable speed?
edit: instead I'll ask this, what's the best model to run on a system with a 24gb 4090 and 64gb of ram?
You can run 4-bit quantized 65B models on your cpu, but it is slow, 1-2 tokens a second instead of 8-15 people typically get with a gpu, but you need two 24gb or an enterprise card with 48gb of ram to load them there.
https://old.reddit.com/r/LocalLLaMA/wiki/models has the information you are irritated about not being listed.
"Reasonable speed" is subjective, though.
A basic question: How can i join with SQL a column to a string separated with comma
GPT-4: PostgreSQL: SELECT STRING_AGG(columnName, ', ') FROM tableName;
Guanaco: Here is an example of how you could use the CONCAT function in MySQL to concatenate a string to a column value in a single-line query: SELECT CONCAT('The total price for ', product_name, ’ is ', SUM(price)) AS total FROM products; This will result in output like “The total price for Chocolate Bar is 10” or similar depending on your data.
Yeah, no...
I think I might give this one to Guanaco. I’m moderately familiar with SQL, and I can’t really understand your question. You’re faking the reader out by using “join” to mean something which is not JOIN.
So do you mean to concatenate a string in a column (coming from a single row) to a “string separated with a comma”? If so (and assuming “separated” means starting or ending with), then Guanaco nailed it. If you meant to join (in the Python sense) the values from a given column in multiple rows, then GPT-4 is doing better.
But GPT-4 gave no explanation, and my general experience with it is that it’s happy to write code that does something vaguely related to the prompt. As the prompt gets more complex or unusual, the degree to which the code doesn’t actually do something useful increases. And I have had quite poor results getting ChatGPT to generate code along with some explanation such that the code actually matches the explanation.
Yeah and not having to do precise prompt engineering is a major advantage of GPT-4 vs other models.
A skilled human will distinguish themselves from a less skilled human by also thinking about the overall problem and helping make sure the specification is correct. This far, I have not seen an AI meet that bar unless the problem is something well-covered already on the web.
I have seen ChatGPT (3.5 and 4) output code that is syntactically correct, executes without errors, and does the wrong things in ways that are not entirely obvious to me until I read the docs myself.
I've only used Bard so far, but indeed, it has never replied to my question with a clarifying question. Only with a statement. This seems like an obvious evolution of the chatbot concept...
edit: I guess GPT-4 has better reading comprension than me!
> How can I join a column with SQL in Postgres to a string separated with a comma?
You can use the `string_agg` function in PostgreSQL to concatenate columns into a string. For example, to concatenate all values from a column named "column1" into a single string separated by commas, you can run the following query:
```sql
SELECT string_agg(column1, ',') FROM table;
```
The `string_agg` function takes two arguments: the first argument is the expression to concatenate (in this case column1), and the second argument is the separator between each value (in this case a comma). The query above will return a single string containing all values from column1 separated by commas.
I generally find WizardLM 1.0, Guanaco and Manticore to be good enough for most GPT cases.
As a human, I wouldn't know what to give you for this either. Not surprised the models give unexpected results. "Guanaco" gives something more in line with what I understand you to be asking for, but clearly you're asking for something else that I don't understand.
To join a column with SQL in Postgres to a string separated by a comma, you can use the STRING_AGG function. Here's an example query that demonstrates this:
SELECT STRING_AGG(column_name, ', ') WITHIN GROUP (ORDER BY id) AS joined_string FROM table;
In this query, we're selecting the joined_string column from a table called table. The STRING_AGG function is used to concatenate all values in the column_name column into a single string separated by commas. We use the WITHIN GROUP (ORDER BY id) clause to specify that we want to group the results by an integer value called id. This ensures that each row of data is treated as a separate entity, even if it has the same column_name.
The resulting output will be a single column with all values from the joined_string column concatenated into a string separated by commas. You can then use this query in your application to join multiple columns together and store them as a single string value.
[1] https://gpt4all.io/reports/GPT4All_Technical_Report_3.pdfThis model seems roughly on par with GPT-3, maybe GPT-3.5 in some cases.
> Write a function to compute the maximum, minimum, and average of an array of doubles in C.
Here is an example function in C that computes the maximum, minimum, and average of an array of doubles:
```c void compute_statistics(double arr[], int n) { double max = arr[0]; double min = arr[0]; double sum = 0.0; for (int i = 1; i < n; i++) { if (arr[i] > max) { max = arr[i]; } else if (arr[i] < min) { min = arr[i]; } sum += arr[i]; } double avg = sum / n; printf("Maximum: %.2f", max); printf("Minimum: %.2f", min); printf("Average: %.2f", avg); } ```
To use this function, simply call it with the array of doubles and its length as arguments. For example:
```c double arr[] = {1.0, 2.5, 3.0, 4.0, 5.0}; int n = sizeof(arr) / sizeof(arr[0]); compute_statistics(arr, n); ```
This will output:
``` Maximum: 5.00 Minimum: 1.00 Average: 3.00 ```
Here are some simple logic questions:
> Compute the sum of an array of 512 elements with each element having the value 2.
The sum of an array of 512 elements with each element having the value 2 is 1024.
> Compute the sum of an array of 512 elements with each element having the value 3.
The sum of an array of 512 elements with each element having the value 3 is 1536.
Not bad.
I just ask "What is the capital of France?" and I haven't had a correct answer yet in any model I tried. They often get Paris right at least, but most other details are wrong.
Guanaco says:
> The current capital of France is Paris. It has been so since 1982 when it replaced the previous one which was Vichy.
My conclusion is that they are using the data from users to improve or some other prompt injection tricks.
So the comparision is not even fair.
Well, you asked for sql and ansi sql function is listagg(), so technically this is not the right answer - and definitely a completely useless one for mysql or oracle users.
MySQL: SELECT GROUP_CONCAT(columnName SEPARATOR ', ') FROM tableName;
Oracle: SELECT LISTAGG(columnName, ', ') WITHIN GROUP (ORDER BY columnName) FROM tableName;
I mean, it's not a big deal.. you just give have to give it a little nudge.
I would assume the poster just edited this though.
Easy way to self-host it is to use text-generation-webui[1] and 33B 4-bit quantized GGML model from TheBloke[2].
I'd really like someone to make a big leaderboard/ranking engine which pits all these engines against eachother and publishes the resulting Elo score.
The precipitating factor is that running large models for research is very expensive, but pales in comparison to putting these things into production. Expenses rise exponentially with model size. Everyone is looking for ways to make the models smaller and run at the edge. I will note that PaLM 2 is smaller than PaLM, the first time I can remember something like that happening. The smallest version of PaLM 2 can run at the edge. Small is beautiful.
Works on all platforms, but runs much better on Linux.
Running this in Docker on my 2080Ti, can barely fit 13B-4bit models into 11G of VRAM, but it works fine, produces around 10-15 tokens/second most of the time. It also has an API, that you can use with something like LangChain.
Supports multiple ways to run the models, purely with CUDA (I think AMD support is coming too) or on CPU with llama.cpp (also possible to offload part of the model to GPU VRAM, but the performance is still nowhere near CUDA).
Don't expect open-source models to perform as well as ChatGPT though, they're still pretty limited in comparison. Good place to get the models is TheBloke's page - https://huggingface.co/TheBloke. Tom converts popular LLM builds into multiple formats that you can use with textgen and he's a pillar of local LLM community.
I'm still learning how to fine-tune/train LoRAs, it's pretty finicky, but promising, I'd like to be able to feed personal data into the model and have it reliably answer questions.
In my opinion, these developments are way more exciting than whatever OpenAI is doing. No way I'm pushing my chatlogs into some corp datacenter, but running locally and storing checkpoints safely would achieve my end-goal of having it "impersonate" myself on the web.
The “best” self-hostable model is a moving target. As of this writing it’s probably one of Vicuña 13B, Wizard 30B, or maybe Guanaco 65B. I’d like to say that Guanaco is wildly better than Vicuña, what with its 5x larger size. But… that seems very task dependent.
As anecdata: my experience is that none of these is as good as even GPT3.5 for summarization, extraction, sentiment analysis, or assistance with writing code. Figuring out how to run them is painful. The speed at which their unquantized variants run on any hardware I have access to is painful. Sorting through licensing is… also painful.
And again: they’re nowhere close to GPT-4.
I use it all the time at home. It's decent at things like summaries and writing content. It follows instructions as well as GPT and isn't nearly as pretentious but it's still a bit pretentious.
It's not very good at code.
https://chat.lmsys.org/?leaderboard
The short answer is that nothing self hosted can come close to GPT-4. The only thing that comes close period is Anthropic's Claude.
There are open source models that are fine tuned for different tasks, and if you're able to pick a specific model for a specific use case you'll get better results.
---
For example, for chat there are models like `mpt-7b-chat` or `GPT4All-13B-snoozy` or `vicuna` that do okay for chat, but are not great at reasoning or code.
Other models are designed for just direct instruction following, but are worse at chat `mpt-7b-instruct`
Meanwhile, there are models designed for code completion like from replit and HuggingFace (`starcoder`) that do decently for programming but not other tasks.
---
For UI the easiest way to get a feel for quality of each of the models (or, chat models at least) is probably https://gpt4all.io/.
And as others have mentioned, for providing an API that's compatible with OpenAI, https://github.com/go-skynet/LocalAI seems to be the frontrunner at the moment.
---
For the project I'm working on (in bio) we're currently struggling with this problem too since we want a nice UI, good performance, and the ability for people to keep their data local.
So at least for the moment, there's no single drop-in replacement for all tasks. But things are changing every week and every day, and I believe that open-source and local can be competitive in the end.
For compatibility with the OpenAI API one project to consider is https://github.com/go-skynet/LocalAI
None of the open models are close to GPT-4 yet, but some of the LLaMA derivatives feel similar to GPT3.5.
Licenses are a big question though: if you want something you can use for commercial purposes your options are much more limited.
I'm the founder of Mirage Studio and we created https://www.mirage-studio.io/private_chatgpt. A privacy-first ChatGPT alternative that can be hosted on-premise or on a leading EU cloud provider.
Wizardlm-uncensored-30B is fun to play with.
(You can use any ChatGPT front-end which lets you change the OpenAI endpoint URL.)
[0] https://huggingface.co/TheBloke/guanaco-65B-HF A QLoRA finetune of LLaMA-65B by Tim Dettmers from the paper here: https://arxiv.org/abs/2305.14314
As for open models, HuggingFace has a nice leaderboard to see which ones are decent: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
Most of the open source stuff people are talking about is things like running a quantized 33B parameter LLaMA model on a 3090. That can be done on consumer hardware, but isn't quite as good at general purpose queries as GPT-4. Depending on your use case and your ability to fine tune it, that might be sufficient for a number of applications. Partcularly if you've got a very specific task.
However, if you're willing to spend, there are bigger models available (e.g. Falcon 40B, LLaMA 65B) that can be run on data server class machines, if you're willing to spend $15-20K.
Will that get you GPT-4 level inference? Probably not (though it is difficult to quantify); will it get you a high-quality model that can be further fine-tuned on your own data? Yes.
For the smaller models, the fine-tunes for various tasks can be fairly effective; in a few more weeks I expect that they'll have continued to improve significantly. There's new capabilities being added every week.
The biggest weakness that's been highlighted in research is that the open source models aren't as good at the wide range of tasks that OpenAI's RLHF has covered; that's partly a data issue and partly a training issue.
For general use Falcon seems to be the current best:
For code specifically Replit's model seems to be the best:
[0]: https://huggingface.co/tiiuae/falcon-40b-instruct [1]: https://huggingface.co/tiiuae/falcon-40b-instruct/blob/main/...
EDIT: I just realized you seem to be asking for a fully realized, turn-key commercial solution. Yeah, refer to others who say there's no alternative. It's true. Something like this gives you a lot more power and flexibility, but at the cost of a lot more work building the solution as you try to apply it.
It's more than that, it requires permission in advance and royalty payments (art. 8). See also attribution requirements in art. 5. Arguably, any license could be waived - it should be possible to write to Meta for permission as well, so it's not in a much better state than Llama itself for commercial use.
I'm especially interested since the data center I'm working for is sitting on a bunch of A100 and I get daily requests of people asking for LLMs tuned to specific cases, who can't or won't use OpenAI for various reasons.
They also have A/B testing with a leaderboard where vicunia wins for the self-hostable ones: https://chat.lmsys.org/?leaderboard
https://lmsys.org/blog/2023-05-25-leaderboard/
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
https://assets-global.website-files.com/61fd4eb76a8d78bc0676...
https://www.mosaicml.com/blog/mpt-7b
Also keep up to date with r/LocalLLaMA where new best open models are posted all the time.
https://lmsys.org/blog/2023-05-25-leaderboard/
But unfortunately for now it seems there aren't any viable self-hosted options...
It's a simple app download and allows you to select from multiple available models. No hacking required.
Guanaco-65B is per https://arxiv.org/abs/2305.14314
CPU Version: https://huggingface.co/TheBloke/guanaco-65B-GGML
GPU Version: https://huggingface.co/TheBloke/guanaco-65B-HF
4bit GPU Version: https://huggingface.co/TheBloke/guanaco-65B-GPTQ
The more you can fit on your GPU (in VRAM) the better (for speed), but no GPU is strictly required.
While this is a MINST classifier (and can be run in the browser) you can get an idea of the math behind it. https://www.3blue1brown.com/lessons/gradient-descent and https://www.3blue1brown.com/lessons/neural-network-analysis
When dealing with a LLM, it's being run again and again - token by token. Pick the best next token, append it to the input, run it again.
If you want to generate 100 tokens (rather small amount of data when compared to much of the GPT-4 conversations), that means running it 100 times.
The network traffic makes the entire system much slower than running it all in one spot.
Consider also the "are you the input to everyone?" and the privacy implications of that.
I think the problem is that it's hard to parallelize this effectively across a network with Internet-scale latency - individual matrix multiplications parallelize very well, but there would need to be coordination of each result, which would be much slower than just doing the computation locally.
In the case of Folding@home, evaluating possible folds could be done completely in parallel and only need to be coordinated on discovery of a plausible fold (which is rare), so distributing over a high-latency network is still beneficial.
I also remember something more directly analogous to @home, but I'm having a hard time finding it.
Edit: Petals! Thank you sibling commenter!
But it is definitely no GPT3.5 or GPT4 replacement. It will not be good for getting "work" done or helping do tasks. It's for recreation. If you want a GPT3.5 level LLM, something akin to text-davinci-002, you'll need to do the SFT and RLHF fine tuning of the 65B llama yourself. And that's no small task. Neither is running a 65B model (even at 4 bits).
GPT AI actually gives me hope. What if we can store and run an AI in a phone-sized-device that is superior to a similarly sized library of books? Can we have a rugged, solar-powered device that could survive the fall of Civilization and help us rebuild?
It would certainly have military applications in a warfare. Imagine being the 21ct century equivalent of a 1940's US Marine on Guadal Canal who need to know some survival skills. ChatGPT-on-a-phone would be handy if you could keep the battery charged.
Would it not make sense to learn survival skills...beforehand?
I would much rather learn survival skills in a controlled environment instead waiting until I need them and then whipping out the ChatGPT enabled phone for help...
Also, I don't need to have an entire library of knowledge with me, when a simple book of survival skills would be more than adequate...
I think pretty much every military in the world agrees with this.
With a 4090, you can get ChatGPT 3.5 level results from Guanaco 33B. Vicuna 13B is a solid performer on more resource-constrained systems.
I'd urge the naysayers who tried the OPT and LLaMA models only to give up to note that the the LLM field is moving very quickly - the current set of models are already vastly superior to the LLaMA models from just two months ago. And there is no sign the progress is slowing - in fact, it seems to be accelerating.
No kidding, and I am calling it on the record right here.
OpenAI will release an 'open source' model to try and recoup their moat in the self hosted / local space.
https://www.theinformation.com/briefings/openai-readies-new-...
Yes it does benefit OpenAI because Sam is pushing for and betting on regulatory capture. Meaning AI models that are compliant with AI safety principles would be allowed and considered safe by law should regulations be put in place.
Current open source models wouldn't be compliant and would require work to make them compliant, thus creating a moat for OpenAI in the self hosted space.
Because there is nothing better than a free model that you can self host that OpenAI has made that is also regulated.
Enterprises would love this, since OpenAI has the mindshare already, they can self host a licensed model in their orgs or air gapped environment that they know is regulated.
OpenAI may release different model sizes of this GPT-X variant like they did quietly for Whisper, Bigger sizes may require an enterprise license.
The big models, if even available, need >100GB of graphics memory to run and would likely take minutes to warm up.
The pricing available via OpenAI/GCP/etc is only effective when you can multi-tenant many users. The cost to run one of these systems for private use would be ~$250k per year.
It's easy to run a much worse model on much worse hardware, but there's a reason why it's only companies with huge datacenter investments running the top models.
It’s actually impressive how good it is considering the limited resources they have.
Of course, running a 180B dense transformer at home for personal use is utterly impractical.