These LLM models have no benefit from running "in the cloud" except for processing power. Lots of disadvantages though, especially in data safety, "leaked chats to other users", privacy, bans etc.
These LLM models have no benefit from running "in the cloud" except for processing power. Lots of disadvantages though, especially in data safety, "leaked chats to other users", privacy, bans etc.
Then there is also moore's law :)
But rn it's totally doable to run a 65B or 100B model on CPU with a reasonable workstation.
Does Moore keep us from wiring more memory onto a GPU? Or making a GPU with expandable memory (slots?)
But there should also be absolutely no issue in making a commercial TPU like google has internally for inference with more but less expensive ram and sell it. There surely must be a market now with these new models.
Now that there's clear demand in the hobbiest market for GPUs >100GB of vram, its more likely that manufacturers will step up with cheaper solutions.
You get fundamentally more powers when you add more VRAM in ways that are just hard to explain to folks outside of this ecosystem. Everything around the VRAM are basically small details in comparison
...with current algorithms and our lack of understanding and insight in to how/why they work on a deep level or what intelligence and consciousness is.
With time hopefully all of these will improve and perhaps future AI's of good quality will be affordable to mere mortals.
A100 seem in the €15k ballpark and H100 double that.
Lot of money but I am actually surprised. A dedicated regular guy could buy this. I mean people buy cars and don’t really need them either. Again not saying it is a bargain, but it’s not billionaires only territory and that is good news (it’s early days!).
The OpenAssistant was/is trained on well structured data from humans for exactly that purpose, for deep learning. In the past most LLMs were trained on unstructured internet data, and they performed well enough. But it was only when OpenAI used reinforcement learning that really the model started to shine.
In my opinion well structured data as input to the machine, have a long way to go. More lightweight models, a lot more precise, a lot faster execution and a lot less memory usage are certainly possible. Most probably we are at the end of the road for the usefulness of structured data. I remember reading an article "Why Large Language models are over", meaning that smaller models but better trained, with better data and algorithms are the way to go.
It feels extremely naive to think that all bans are a bad thing.
Let's say that a criminal org starts a fully automated system to scam grandmas out of their savings. A cloud based service could ban them. A self-hosted system could not.
Yet it is widely regarded as a good thing that nmap can be distributed and printers can be bought. Why are these models special?
Whereas nmap and printers cannot (yet)
Realistically though and as we have seen with ChatGPT, if models can be censored they will be censored to the point where it affects normal people. Most people using chatbots have experienced "as an AI model, I can't do that" because of bullshit ethics.
So... sorry about your savings Grandma, but I'm still going to fight for uncensored AI models. Fraud is already illegal, and if it happens we can prosecute the offenders.
Grandma isn't leaving you anything when she passes if all of her savings were plundered by scammers while she was still alive.
Preventing elder abuse is in your best interest.