Why host your own LLM?
marble.onl
marble.onl
Latency is down consistently across the board and we haven't seen a single "429 model overloaded" error in the past month: https://twitter.com/_cartermp/status/1686894576202907651/
Though GPU workloads are still a point where building your own server and running it from your basement or putting it in colo can be very attractive.
from beam import App, Runtime
app = App(name="gpu-app", runtime=Runtime(gpu="T4"))
@app.rest_api()
def inference():
print("This is running on a GPU")
Then run beam deploy {your-app}.py and boom, it's running on the cloudIf it works the way I think it does, it sounds appealing, but the GPUs also feel a bit small. The A10G only has 24GB of VRAM. They say they're planning to add an A100 option, but... only the 40GB model? Nvidia has offered an 80GB A100 for several years now, which seems like it would be far more useful for pushing the limits of today's 70B+ parameter models. Quantization can get a 70B parameter model running on less VRAM, but it's definitely a trade-off, and I'm not sure how the training side of things works with regards to quantized models.
Beam's focus on Python apps makes a lot of sense, but what if I want to run `llama.cpp`?
Anyways, Beam is obviously a very small team, so they can't solve every problem for every person.
[0]: what is the "time to idle" for serverless functions? is it instant? "Pay for what you use, down to the second" sounds good in theory, but AWS also uses per-second billing on tons of stuff... EC2 instances don't just stop billing you when they go idle, though, you have to manually shut them down and start them up. So, making the lifecycle clearer would be great. Even a quick example of how you would be billed might be helpful.
An A4000 for 139$ x 12 is not terrible
Tiny Corp. will be producing the Tiny Box that lets you host your own LLM at home using TinyGrad software.
The tinybox
738 FP16 TFLOPS
144 GB GPU RAM
5.76 TB/s RAM bandwidth
30 GB/s model load bandwidth (big llama loads in around 4 seconds)
AMD EPYC CPU
1600W (one 120V outlet)
Runs 65B FP16 LLaMA out of the box (using tinygrad, subject to software development risks)
$15,000
Hardware startups are extremely difficult. But Hotz's other company Comma.ai is already profitable so it is possible. I find the guy extremely encouraging and he is always doing interesting stuff.
(If money is no object, why not grab an oxide.computer rack? Assuming you have three-phase power, of course...)
At work I've been successfully running llama2 variants on Mac M2 systems with more than adequate throughput for small experiments and examples. Can also boot a 4x Nvidia T4 instance on AWS (g4dn.12xlarge) to larger models.
I feel you'd need to be doing inference on a 24/7 basis for at least a year to justify the price of this machine.
> three-phase power,
Think for the tiny grad it's run off single, but for the US the GPUs are power limited given the 110V grid.
For something that is... going to hallucinate? I don't get it?
With this, it could always be repurposed if something new and shiny comes along
I mean, the oxide.computer specs make me go "this would be the most killer gaming machine in all of existence" ... though it's a server rack meant to host an entire private cloud for an entire corporation, I still would love to glom together as much resources as can fit into a single VM for gaming purposes.
That and setting up the software to work w/ an array of GPUs rather than one seems difficult.
It probably depends on your gpu though.
Undervolting would probably be even better.
~ > sudo nvidia-smi -pl 300
Power limit for GPU 00000000:09:00.0 was set to 300.00 W from 450.00 W.Was it this one? https://www.youtube.com/watch?v=dNrTrx42DGQ
But isn't M2 ULTRA over 20x slower than this thing? ~30 TFlops vs 738.
1. We are not permitted, by rule, to send client data to unapproved third parties.
2. I generally do not trust third parties with our data, even if it falls outside of #1. Just look at the hoopla with Zoom; do you really want OpenAI further solidifying their grip on the industry with your data?
3. We have the opportunity to refine the models ourselves to get better results than offered out of the box.
4. It's fun work to do and there's a confluence of a ton of new and existing nerdy technology to learn and use.
I'm very curious what your experience is? Do you think to self-host is good enough so users will accept it?
2. We tested 13 and 30b parameter pretrained LLMs and found them “good enough” within the confines of the original scope.
That said, we are expanding the use cases and with that comes the need for increasing levels of quality. This is where the self-training comes in.
We’re fairly confident it’ll meet our needs for now.
EDIT: Or in the worst case ignore the policy and get fired.
We will see how this turns out with LLMs and while I wish companies were more concerned about what data they share with third parties my experience is that apart from military/intelligence related data everything can and will be outsourced.
I have this urge to toy with the idea but i also find "Prompt Engineering" to be very unattractive. It feels like something i'd have to re-tailor towards any new model i change to. Not very re-usable and difficult to make "just work".
Mind elaborating on that? Looks like a typo but i'm having difficulty knowing for sure. Thanks!
It's not perfect but it removed a lot of useless stuff like all the times i mistyped '--recurse-submodules' or the times I went into a folder only to realize it was not the correct folder for what I was trying to do.
What would be your goal for doing that?
Sounds a bit silly, but of course it's mostly just for fun. However on the more practical side, i do often find myself needing to dig through old text conversations trying to find that one message. Not having a flexible, deep search behind it sucks. I often find myself wanting to do the same with my browser history. Find that one website i visited, etc.
I have the thought that it would be great to make my data points more rich. Don't just tag my browser history with isolated tags, such as Programming, Rust, etc - but infer meaning from my searching. Be able to see that i'm working on ProjectX actively via CLI Git activity, and that i'm searching for Y. Be able to correlate commit Z with search Y. etcetc
It feels to me there's a ton of small, edge case utility that can be gained by dumping everything to a local server and having it link the data. But i don't want to do any of that manually.
Likewise, i've wanted to manage "Home Inventory" before - what's in boxes, etc. Managing that myself is tedious, though. LLMs seem ripe for figuring out associations - even dumb LLMs. My hope is that i eventually can start wiring things together and having the LLMs start making rich data out of messy untagged data.
Would be neat, /shrug
Personally i want to build a fairly dumb system though. Ie make a system which can be useful with LLama2 13B or w/e. Something that doesn't require state of the art GPT4+.
If that means compromising on some features that's fine, but at least then it can be truly and fully local.
What LLMs are you using?
> how is the performance/cost of running vs more specialized models trained for the task. most models are GNU licensed so thats not an issue. But I imagine you meant the age old question of hosting yourself vs using openAI. Truth is as of now it currently is not foretasted to beat using one of the less intelligent models on openAI. hardware cost alone yes but Dev time is very expensive. Lucky were a small company & our CEO sees this as training. Because LLMs are so new there really isn't a large labor market for it yet. If our devs and engineers get in this early then we can beat others to market as the technology develops and new opportunities come to light. on top of having possible HIPPA, GDPR, or other security laws to follow that OpenAI has been very shooty about, we do not want be at the whim of OpenAI or another SaaS provider on a mission critical part of a vertical. They have talked about depreciating old model. As well they have had content changes in there models to placate political critics, well not realizing that this pulls the rug out from under developers that need any sense of stability from there product.
But the dev and computing cost of this feels so huge that I'm not even sure where to start.
While it really resonated with my management I felt worried I wouldn't be able to replicate these kind of results on other projects.
THE ONLY REAL ADVICE I CAN GIVE ON AI PROJECTS IS . . . don't let your managements expectation of LLMs out weigh its capabilities.
I'm sure I speak for many people here when your non-tech fluent directors get together and think GPT4 is some sort of deity. GPT4 smart (or used to be at least) ill give it that, but small locally hosted 7b/13b LLMs are very limited and people for whatever reason get AI infatuation the second they finally see you show direct value in it they will lose there shit in its assumed capabilities. you got to be direct with them that no matter what dumb video they saw on Sam Altman, what your are proposing is not that. Be very clear in its possible scope because there is some idiot in our organization that will assume assume you can programmatically answer prayers. I actually had this guy from our networking team try and raise a concern about the LLM going sentient and us having a "Skynet" problem. granted this was back in march/2023 so AI histira was a little more rampant but still.
tl;dr my recommendation for your pdf project is run https://github.com/oobabooga/text-generation-webui. if your can get a 30 series GPU in your company Then run a 13B 4bit model that can pull info, assign tags, run minor analysis on your text. else find a spare 16gb machine and do the same but but over a longer time scale.
run a prompt that checks for hallucinations. "does the following text make sense? previous prompt + text if yes then keep else make intern do it.
GPT-j-7b is still one of the best models because it has indexing & categorizing at the main prosperous. other models are great but core idea behind LLMs is that its just a high level auto complete
By fine-tuning the model to extract a specific desired output from the text you give it, it learns that the output always comes from the input, and so you get less random outputs than just by prompting an instruction-tuned model (which was fine-tuned to find the answer in its weights, instead of copying it from the input).
It seems like llama2 is the biggest name on HN when it comes to self hosting but I have no idea how it actually performs.
Grab KoboldCPP and a GGML model from TheBloke that would fit your RAM/VRAM and try it.
Make sure you follow the prompt structure for the model that you will see on TheBloke's download page for the model (very important).
KoboldCPP: https://github.com/LostRuins/koboldcpp
TheBloke: https://huggingface.co/TheBloke
I would start with a 13b or 7b model quantized to 4-bits just to get the hang of it. Some generic or story telling model.
Just make sure you follow the prompt structure that the model card lists.
KoboldCPP is very easy to use. You just drag the model file onto the executable, wait till it loads and go to the web interface.
Ie how do you feed the LLM the text along with your question without it forgetting most of the text? I assume the text you want to feed it is longer than 16,000 words.
How do you do the tagging bits, though?
If you have a product that uses an LLM and can get away with one of the open source ones, it’s probably cheaper (and def lower latency/response time) to host yourself too somewhere like azure or aws.
I run oobabooga's API on docker with a 13B 4bit quantized model. https://github.com/oobabooga/text-generation-webui
We use GTX 3060s because there the best bang for there buck in terms of VRAM. Our current set up is mostly proof of concept or used of inner office work well we work on scaling to get a fluid handler built so it can distribute workloads around the multiple GPUs.
Lucky the crypto mining community laid the ground work for some of the hardware.
Are you using riser cards to connect the GPU's to the motherboards then? I thought about trying a setup like yours, but was worried that the riser card interfaces would create a bottleneck. Ideally I'd like to run some cards in a separate box and connect them to my main computer through some kind of cable interface, but I'm not sure if that's possible without seriously affecting performance.
8-bit quantized uncensored Llama 2 13B, doing 50 generated tokens/second, using CPU+GPU including 17GB of 3090's 24GB VRAM.
I also have quantized 70B running currently CPU-only, but I might later be able to speed that up with some CUDA or OpenCL offloading.
This is on Debian Stable (like usual), albeit currently with closed Nvidia CUDA stack, and necessarily with the closed Llama 2 that I can only fine-tune atop. (I'm hoping that some scientific/academic non-profit/govt effort will be able to muster fully open models in the future.)
One of the main reasons I picked Llama 2 was the relatively friendly licensing (and Meta is earning lots of goodwill with that). With this licensing, and the performance I'm getting, in theory, I could even shoestring bootstrap an indie startup with low online LLM demands, from a single consumer hardware box in the proverbial startup garage or kitchen table. (Though I'd try to get affordable cloud compute first.)
One of our big questions is whether it makes sense to rent or to buy for training/finetuning/RLHF. The advantage of renting is obvious: I don't think that this phase of the project will last very long, and if it turns out that the idea is a success we'll have no problem securing funding for perma-improvement infra.
The possible advantage of buying is that we would then have the hardware available for inference hosting. We do expect some amount of demand in perpetuity. Having that ongoing cost as small as possible would allow us to continue serving the "clients" we KNOW would benefit a lot from our service with minimal recurring revenue.
70B is currently 4-bit on this box, and once I have GPU accel for 70B, I'll see how the quality compares to 13B 8-bit.
https://huggingface.co/Open-Orca/OpenOrca-Platypus2-13B
https://huggingface.co/spaces/Open-Orca/OpenOrca-Platypus2-1...
A few more:
- Wizard Vicuna 13B uncensored
- Nous Hermes Llama 2
- WizardLM Uncensored llama2
I'm only waiting to have decent LLMs able locally to be able to start really using them.
No way I'd feed my code, customer data, personal info, secrets, emails, etc., to some dubious cloud machine which is already excellent at exctracting valuable or juicy bits from what I'm feeding it.
I have a little quip I use to troll people who say in effect- but but contract law, your contract says your data is yours even when it is on someone's cloud servers- I say- I know, I know, but remember "possession is 9/10ths of the law."
I do believe in 99.(many 9s) cases no one admin with visibility cares about any given customer's stuff but if they do and if it matters, by then it's too late.
- Corsair H150i PRO 47.3 CFM Liquid CPU Cooler
- MSI MPG Z590 GAMING EDGE WIFI ATX LGA1200 Motherboard
- G.Skill Ripjaws V 64 GB (4 x 16 GB) DDR4-3600 CL18 Memory
- Samsung 970 Evo Plus 1 TB M.2-2280 PCIe 3.0 X4 NVME Solid State Drive
- MSI GeForce RTX 3090 TI SUPRIM X 24G GeForce RTX 3090 Ti 24 GB Video Card
- Corsair Carbide Series 275R ATX Mid Tower Case
- Corsair RM1000x (2021) 1000 W 80+ Gold Certified Fully Modular ATX Power Supply
- Microsoft Windows 11 Pro
On this setup I've been able to run every model 13B and below with 0 issue. Even been able to fine-tune Llama 2 13B using my own data (emails, SMS, FB messages, WhatsApp, etc.) with pretty fun results!
Is it useful? Would you let it draft responses at you? Curious about the fun results :)
Find a repo. follow the install instructions. What is this weird error? A library issue..? Maybe it's my OS..?
It always seems to be tedious compared to open projects in other domains. Maybe that can't be solved.
Honestly the easiest way that “just works” is to use LM Studio which you can run locally https://lmstudio.ai/
Obviously you’ll have faster results if you have a fancy gaming GPU or something like the M2 Max/Ultra but you don’t need those to have a play and see if it interests you.
[0] https://gpt4all.io/index.html [1] https://docs.gpt4all.io/
Using docker has an initial threshold you have to get over but once you do, everything becomes very easy in it. How you end up using docker matters very little once you get the concepts.
- I use LLMs for code generation for a startup and they are not competitive for that yet.
- Most of the popular open models are non-commercial.
- The only practical way I know of to get large custom datasets for training is to have OpenAI's models generate them, and they forbid this in their terms of service.
Having something that's truly open and closer to GPT-4 for code generation will probably happen within less than a year (I hope) and will be a game changer for self-hosting.
I have not tried public LLMs myself. Do they give reproducible results?
Public LLMs I don't know but images generated using StableDiffusion are, of course, totally deterministic.
There really is no reason a LLM cannot be deterministic and if it isn't: fix it (even if this comes at a tiny performance cost).
Sometimes answer is wrong and then right.
If it is deterministic, then what if gets "stuck" on wrong answer?