Cost of self hosting Llama-3 8B-Instruct
blog.lytix.co
blog.lytix.co
Assuming we want to mirror our setup in AWS, we’d need 4x NVidia Tesla T4s. You can buy them for about $700 on eBay.
Add in $1,000 to setup the rest of the rig and you have a final price of around:
$2,800 + $1,000 = $3,800
This whole exercise assumes that you're using the Llama 3 8b model. At full fp16 precision that will fit in one 3090 or 4090 GPU (the int8 version will too, and run faster, with very little degradation.) Especially if you're willing to buy GPU hardware from eBay, that will cost significantly less.
I have my home workstation with a 4090 exposed as a vLLM service to an AWS environment where I access it via reverse SSH tunnel.
I'm guessing actual break-even time is less than half that, so maybe 2 years.
This whole post makes OpenAI look like a better deal than it actually is.
EDIT: TOS -> EULA per comments below
Regardless, it is true that it is a problem.
https://www.reddit.com/r/MachineLearning/comments/ikrk4u/d_c...
What if I’m a business and I’m selling LLMs in a box for you to put on a private network?
What constitutes a data center according to the ToS? Is it enforceable if you never agree to the ToS (buying through eBay?)
> No Datacenter Deployment. The SOFTWARE is not licensed for datacenter deployment, except that blockchain processing in a datacenter is permitted.
- https://www.nvidia.com/content/DriverDownloads/licence.php?l...
There's no further elaboration on what "datacenter" means here, and it's a fair argument to say that a closet with one consumer-GPU-enriched PC is not a "datacenter deployment". The odds that Nvidia would pursue a claim against an individual or small business who used it that way is infinitesimal.
So both the ethical issue (it's a fair-if-debatable read of the clause) and the practical legal issue (Nvidia wouldn't bother to argue either way) seem to say one needn't worry about.
The clause is there to deter at-scale commercial service providers from buying up the consumer card market.
No one cares about this TOS provision. I know both startups and large businesses that violate it as well as industry datacenters and academic clusters. There are companies that explicitly sell you hardware to violate it. Heck, Nvidia will even give you a discount when you buy the hardware to violate it in large enough volume!
You do you.
If you don't rent our servers or VMs Nvidia doesn't care. They aren't Oracle.
I would not bank on anything in this article. It might as well have been written by a tiny Llama model.
Have a 13B running on a 3070 with 16 gpu layers and the rest running off CPU.
Performs okay, but way cheaper than renting a GPU on the cloud.
brew install ollama; ollama serve; ollama pull llama3: 8b-v2.9-q5_K_M; ollama run llama3: 8b-v2.9-q5_K_M
https://ollama.com/library/dolphin-llama3:8b-v2.9-q5_K_M
(It may need to be Q4 or Q3 instead of Q5 depending on how the RAM shakes out. But the Q5_K_M quantization (k-quantization is the term) is generally the best balance of size vs performance vs intelligence if you can run it, followed by Q4_K_M. Running Q6, Q8, or fp16 is of course even better but you’re nowhere near fitting that on 8gb.)
https://old.reddit.com/r/LocalLLaMA/comments/1ba55rj/overvie...
Dolphin-llama3 is generally more compliant and I’d recommend that over just the base model. It's been fine-tuned to filter out the dumb "sorry I can't do that" battle, and it turns out this also increases the quality of the results (by limiting the space you're generating, you also limit the quality of the results).
https://erichartford.com/uncensored-models
https://arxiv.org/abs/2308.13449
Most of the time you will want to look for an "instruct" model, if it doesn't have the instruct suffix it'll normally be a "fill in the blank" model that finishes what it thinks is the pattern in the input, rather than generate a textual answer to a question. But ollama typically pulls the instruct models into their repos.
(sometimes you will see this even with instruct models, especially if they're misconfigured. When llama3 non-dolphin first came out I played with it and I'd get answers that looked like stackoverflow format or quora format responses with ""scores"" etc, either as the full output or mixed in. Presumably a misconfigured model, or they pulled in a non-instruct model, or something.)
Dolphin-mixtral:8x7b-v2.7 is where things get really interesting imo. I have 64gb and 32gb machines and so far the Q6 and q4-k_m are the best options for those machines. dolphin-llama3 is reasonable but dolphin-mixtral is a richer better response.
I’m told there’s better stuff available now, but not sure what a good choice would be for for 64gb and 32gb if not mixtral.
Also, just keep an eye on r/LocalLLaMA in general, that's where all the enthusiasts hang out.
Just download a single file and run it.
It’s been an absolute gamechanger.
the docs are good. when creating the initial CA make absolutely sure you set the CA expiration to 10-30 years, the default is 1 which means your whole setup explodes in a year without warning.
It’s end to end encrypted, and with tail lock enabled, nodes can not be added without user’s permission.
I didn’t start with tailscale because the only way you could log into it was with Google or GitHub or something. I don’t trust Microsoft or Google with auth for my internal network. I thought about running Headscale but Nebula was faster/easier for me.
There's also potential for malicious updates to compromise a network (as there is with most software unless you're auditing the source for each update).
E2EE is only as meaningful as where the keys reside, and how easily those keys are abused.
The metadata is generally public information, I don’t care about that.
The malicious updates and key abuse are more concerning. It’s true for all software, and probably better done with OS, like on iOS.
The VPN could steal the keys, but that’s a lawsuit!
But with a malicious update, they could ship them to their infra, targeting some users. The product then becomes malware!
totally different use case.
I think port forwarding configuration is a pain that does not offer value over just poking a hole in your firewall to do an authenticated connection over ssh.
I’ve gone through hundreds of millions, maybe billions, of tokens in the past year.
This article is just “cloud is expensive” 101. Nothing new.
I don't do any work with the output, I'm just the MLOps guy (ahem, DevOps).
You mention expense but on a purely financial basis I find any of these hosted solutions really hard to justify against GPT 3.5 turbo prices, including building your own rig. $5k + electricity is loads of 3.5 Turbo tokens.
Of course none of the data scientists or researchers I work with want to use that though - it's not their job to host these things or worry about the costs.
Knowing up front this is my fixed ML budget gives me peace of mind and gives me room to try stupid ideas without worrying about it.
Whereas doing it in the cloud you can a) get slammed with some crazy bill by accident, b) have to think more about what resources testing an idea will take, or conversely c) getting GPU FOMO and thinking “if just upgrade a level all my problems will be solved”.
It works for me, everybody mileage varies but personally I like to budget; spend; and then totally focus on my goals and not my cloud spend.
I’m also from the pre-cloud era, so doing stuff on my own bare metal is second nature.
I need to tell if I change I made was impactful, but if the model just magically gets smarter or dumber at my tasks with no warning then I can’t tell if I made an improvement or a regression.
Whereas the model on my GPU doesn’t change unless I change it. So it’s one less variable and LLM are black box to start with.
I may be wrong for Gemini, but my impression is all the companies are constantly tweaking the big models. I know GPT on Monday is not always the same GPT on Thursday for example.
Ollama makes it dead simple just to try it out. I was pleasantly surprised by the tokens/sec I could get with Llama 3 8B on a 2021 M1 MBP. Now need to try on my gaming PC I never use. Would be super cool to just have a LLM server on my local network for me and the fam. Exciting times.
LLAMA 8B on Bedrock is $0.40 per 1M input tokens and $0.60 per 1M output tokens which is a lot cheaper than OpenAI models.
Edit: to add to that, as technical people we tend to discount the value of our own time. Bedrock and the OpenAI are both very easy to integrate with and get started. How long did this server take to build? How much time does it take to maintain and make sure all the security patches are applied each month? How often does it crash and how much time will be needed to recover it? Do you keep spare parts on hand / how much is the cost of downtime if you have to wait to get a replacement part in the mail? That's got to be part of the break-even equation.
Either way, I wish Groq would offer a real service to people willing to pay.
I don't really get it, only thing I can surmise is it'd be such a no-brainer in various cases, that if they tried supporting it as a service, they'd have to cut users. I've seen multiple big media company employees begging for some sort of response on their discord.
https://aws.amazon.com/sagemaker/pricing/
I'd love to hear someones experience using this as I want to make an RPG rules bot tied to a specific ruleset as a project but I fear AWS as it might bankrupt me!
To help control costs, you can choose pretty conservative settings in terms of how long you want to let the model train for. Once that iteration is done and you have a model artifact saved, you can always pick back up and perform more rounds of training using the previous checkpoint as a starting point.
About 3 days [from 0 and iterating multiple times to the final solution]
> How much time does it take to maintain and make sure all the security patches are applied each month?
A lot
> How often does it crash and how much time will be needed to recover it? Do you keep spare parts on hand / how much is the cost of downtime if you have to wait to get a replacement part in the mail?
All really good points, the exercise to self host is really just to see what is possible but completely agree that self hosting makes little to no sense unless you have a business case that can justify it.
Not to mention if you sign customers with SLAs and then end up having downtime would put even more pressure on your self hosted hardware
All of these are the kinds of things that people say to non-technical people to try to sell cloud. It's all fluff.
Do you really think that cloud computing doesn't have security issues, or crashes, or data loss, or that it doesn't involve lots of administration? Thinking that we don't know any better is both disingenuous and a bit insulting.
Tbh, your comment is kind of insulting and belittles how far we've come ahead in infrastructure management.
The cloud is probably more secure than a set of janky servers that you have running in your basement. You can totally automate away 0-days, cves and get access to better security primitives.
However, you've now gone out of your way to try to be insulting. You know nothing about me, yet you want to suggest that the cloud is more secure than my servers, and that my servers are "janky"?
Please try a little harder to engage in reasonable discourse.
Apples/Oranges. Your janky cloud[0] is less secure than the servers in my basement, because I'm a mostly competent sysadmin. Cloud lets you trade some operational concerns for higher costs, but not all of them.
[0] If you can assume servers run by somebody who doesn't know how to do it properly, obviously I can assume the same about cloud configuration. Have fun with your leaked API keys.
> doesn't have security issues
It sure does, but the matrix of responsibility is very different when it is a hosted service. Note: I am making these comments about Bedrock, which is serverless not in relation to EC2.
> It crashes
Absolutely, but the recovery profile is not even close to the same. Unless you have a full time person with physical access to your server who can go press buttons.
> data loss
I'm going to shift this one a tiny bit. What about hardware loss? You need backups regardless. On the cloud when a HDD dies you provision a new one. On premise you need to have the replacement there and ready to swap out (unless you want to wait for shipping). Same with all the other components. So you basically need to buy two of everything. If you have a fleet of servers that's not too bad since presumably they aren't going to all fail on the same component at the same time. But for a single server it is literally double the cost.
> doesn't involve lots of administration
Again, this is relation to Bedrock with is a managed serverless environment. So there is litterally no administration aside from provisioning and securing access to the resource. You'd have a point if this was running on EC2 or EKS but that's not what my post was about.
> Thinking that we don't know any better is both disingenuous and a bit insulting.
I'm not saying cloud is perfect in any way, like all things it requires tradeoff, but quite frankly I find you dismissing my 25 years of experience, 1/3 of that has been working in real data centers (including a top-50 internet company at the time) as "fluff" as "disingenuous and a bit insulting".
Every colo facility I’ve used offers “remote hands”. If you need a button pressed or a disk swapped, they will do it, with a fee structure and response time that varies depending on one’s arrangement with the operator. But it’s generally both inexpensive and fast.
> What about hardware loss? You need backups regardless. On the cloud when a HDD dies you provision a new one. On premise you need to have the replacement there and ready to swap out (unless you want to wait for shipping).
Two of everything may still be cheaper than expensive cloud services. But there’s an obvious middle ground: a service contract that guarantees you spare parts and a technician with a designated amount of notice. This service is widely available and reasonably priced. (Don’t believe the listed total prices on the web sites of big name server vendors — they negotiate substantial discounts, even in small quantities.)
I guess inexpensive is relative. I've been on cloud for a while so I'm not sure what the going rates are for "remote hands" and most of my experience is with on-premise vs co-lo.
> Two of everything may still be cheaper than expensive cloud services.
That is true. Everything has tradeoffs. Though in the OPs case I think the math is relatively clear. With Open AIs pricing he calculated the break even at 5 years just for the hardware and electricity. Assuming that calculation is right, two of everything would up that to 7+ years, at which point... a lot can happen in 7 years.
At a low end facility, I’ve usually paid between $0 and $50 per remote hands incident. The staff was friendly and competent, and I had no complaints. The price list goes a bit higher, but I haven’t needed those services at that facility.
And do you really think you can offer better security and uptime than AWS? Not impossible but very expensive if you’re managing everything from your own hardware. You clearly vastly underestimate all that AWS is taking care of.
Overlooking that, the rest of the article feels a bit strange. Would we really have a use case where we can make use of those 157 million tokens a month? Would we really round $50 of energy cost to $100 a month? (Granted, the author didn't include power for the computer) If we buy our own system to run, why would we need to "scale your own hardware"?
I get that this is just to give us an idea of what running something yourself would cost when comparing with services like ChatGPT, but if so, we wouldn't be making most of the choices made here such as getting four NVIDIA Tesla T4 cards.
Memory is cheap, so running Llama-3 entirely on CPU is also an option. It's slower, of course, but it's infinitely more flexible. If I really wanted to spend a lot of time tinkering with LLMs, I'd definitely do this to figure out what I want to run before deciding on GPU hardware, then I'd get GPU hardware that best matches that, instead of the other way around.
No. I googled "self hosting", read the first few definitions, and they agree with the article, not you. E.g., wikipedia -- https://en.wikipedia.org/wiki/Self-hosting_(web_services)
> Self-hosting is the practice of running and maintaining a website or service using a private web server, instead of using a service outside of someone's own control.
Hosting anything on Amazon is not "using a private web server" and is the very definition of using "a service outside of someone's own control".
The fact that the rest of the article talks about "enabled users to run their own servers on remote hardware or virtual machines" is just wrong. It's not "their own servers", and we don't have "more control over their data, privacy" when it's literally in the possession of others.
> The practice of self-hosting web services became more feasible with the development of cloud computing and virtualization technologies, which enabled users to run their own servers on remote hardware or virtual machines. The first public cloud service, Amazon Web Services (AWS), was launched in 2006, offering Simple Storage Service (S3) and Elastic Compute Cloud (EC2) as its initial products.[3]
The mystery deepens
I think we generally understand Iaas, Paas, and Saas, to be hosted offerings, managed and unmanaged...
On demand pricing put it at about $0.45 per million tokens.
Source: We use TPUs at scale at https://osmos.io
Google Next 2024 session going into detail: https://www.youtube.com/watch?v=5QsM1K9ahtw
At some point everyone will realize that it is becoming a commodity, and it is very expensive to train, then only those with wither a structural advantage to lower price (eg Google) or a true goal of being on the high-end/SOTA of the market (OpenAI, Anthropic) will keep going.
I'm finding it difficult to estimate the size of this workload compared to continued training of foundation models. Perhaps it depends on whether there are new architectural breakthroughs that require retraining of foundation models.
And what about non-language tasks such as interpreting video and 3D sensory data? This is potentially huge, but between huge peaks there is often a valley of unknowable depth and breadth.
Yea yes video probably requires a lot of GPUs to train. And a lot of source material to train against. And a use case. Which again, most companies don’t have.
Model development is clearly here to stay, and clearly valuable. Models from every other company, either foundation or fine tuned - I’m not sure that emperor is wearing many clothes any time soon.
There are probably some situations where it suffices to use a small model, but for most purposes, I'd prefer to use the state of the art, and I'm eager for that state to progress a little more.
I'm guessing in the future, we'll see a lot more automatic inference "on our behalf", and you won't care or notice if its using Llama5:3b or whatever comes out then.
I'm betting that in a few years we'll see LLMs baked into a ton of stuff - lots of "simple" things mostly, like email summary/rewording, and similar "light touch" use cases. Maybe light multi-modal work like photo labeling. Probably expanded to be running in every other SaaS applications doing god knows what. For those, the small models would be more than enough, and much cheaper to run in large volumes.
I'm guessing "chat with a bot to ask questions" will be a small amount of the inference that happens on our behalf, but will use the valuable SOTA model use case.
It’s a wee bit early to call this. Let’s see what the top labs release in the next year or two, yeah?
GPT-4 was released only 15 months ago, which was about 3 years after GPT-3 was released.
These things don’t happen overnight, and many multi-year efforts are currently in the works, especially starting last year.
Not only big tech is part of it but billion dollars startups are popping everywhere from China to US and Middle East.
next wave driving demand can be actual new products developed on LLMs. There are very few usecases currently well developed besides chatbots, but potential is very large.
Your self-hosted LLM box is going to use maybe 20-30% of the power this article suggests it will.
Source: I run LLMs at home on a machine I built myself.
That being said, the difference between OpenAI and AWS cost ($1 vs $17) is huge. Is OpenAI just operating at a massive loss?
Edit: Turns out AWS is actually cheaper if you don't use the terrible setup in this article, see comments below.
A100 TDP is 400W so assuming 4kW for the whole machine, that's a little more than $5k/year at $0.15/kWh. Again, the difference is in the tens of thousands per instance. Even at 50% utilization over three years, if you need more than a dozen machines it's much cheaper to buy them outright, especially on credit.
If you keep reading past there, they get it down significantly. The 8 tkn/s number AWS was evaluated on is really funny, that's about what you'd get on last year's iphone and it's not because apples special, it's because theres barely any reasonable optimization being done here. No batching, float32 weights (8 bit is guaranteed indistinguishable from 32 bit, 5 bit tests as definitely indistinguishable in blind tests, 4 bit arguably is indistinguishable)
So in reality AWS is cheaper for a much better model if you don't go with a wildly suboptimal setup.
even with the subs and api charges, they still let people use chatGPT for free with no monetization options. Sure they are collecting the data for training, but that's hard to quantify the value of.
It is miraculous that the cost comparison isn't worse given how adversarial this test is.
Larger requests, concurrent requests, and request queueing will drastically reduce cost here.
Not following?
Llama 8B is like 17ish gigs. You can throw that onto a single 3090 off ebay. 700 for the card and another 500 for some 2nd hand basic gaming rig.
Plus you don't need a 4 slot PCIE mobo. Plus it's a gen4 pcie card (vs gen3). Plus skipping the complexity of multi-GPU. And wouldn't be surprised if it ends up faster too (everything in one GPU tends to be much faster in my experience, plus 3090 is just organically faster 1:1)
Or if you're feeling extra spicy you can do same on a 7900XTX (inference works fine on those & it's likely that there will be big optimisation gains in next months).
Someone correct me if I'm wrong, but I've always thought you needed enough VRAM to have at least double the model size so that the GPU has enough VRAM for the calculated values from the model. So that 17 GB model requires 34 GB of RAM.
Though you can quantize to fp8/int8 with surprisingly little negative effect and then run that 17 GB model with 17 GB of VRAM.
Here is a calculator (if you have a GPU you want to use EXL2, otherwise GGUF) https://huggingface.co/spaces/NyxKrage/LLM-Model-VRAM-Calcul...
Also model quantisation goes a long way with surprisingly little loss in quality.
Edit: Oh I see that calculator you linked shows that too. My information was more trial and error, thanks for the calculator link.
I think this worked out to around $1/million tokens of output and orders of magnitude less for input tokens, and before reserved instances or other providers were considered.
I sometimes use Groq Llama3 APIs (so fast!) or OpenAI APIs, but I mostly use my 32G M2 system.
The article calculates cost of self-hosting, but I think it is also good taking into account how happy I am self hosting on my own hardware.
To run llama-3 8b. A new $300 3060 12gb will do, it will load fine in Q8 gguf. If you must load in fp16 and cash is a problem a $160 P40 will do. If performance is desired a used 3090 for ~$650 will do.
edit: Actually could you share how long it took to make a query? One of our issues is we need it to respond in a fast time frame
(Haven't used it in production, thinking to use it for side projects).
The pricing is outdated now.
Here is the piece -https://www.inferless.com/learn/unraveling-gpu-inference-cos...
1. Don't use AWS, because it's one of the most expensive cloud providers
2. Use quantized models, because they offer the best output quality per money spent, regardless of the budget
This article, on the other hand, focuses exclusively on running an unquantized model on AWS...
Check out the host LLM's at home crowd. One app to look at is llama.cpp. Model compression is one of the first techniques to successfully run models on low capacity hardware.
>( 100 / 157,075,200 ) * 1,000,000 = $0.000000636637738
Should be $0.64 so still expensive
groq costs about that for llama 3 70b (which is a monumentally better model) and 1/10th of that for llama 3 8b
That said, a while ago there was this[1] thread on here which helped me snatch a brand new, unboxed p40 for peanuts. Really, the cost was 2 or 3 jars of good quality peanut butter. Sadly it's still collecting dust since although my workstation can accommodate it, cooling is a bit of an issue - I 3D printed a bunch of hacky vents but I haven't had the time to put it all together.
The reason why I went this road was phi-3, which blew me away by how powerful, yet compact it is. Again, I would not trust it with anything big, but I have been using it for sifting through a bunch of raw, unstructured text and extract data from it and it's honestly done wonders. Overall, depending on your budget and your goal, running an llm in your home lab is a very appealing idea.
I would like to compare the costs vs hardware on prem, so this helps with one side of the equation
AWS is stupidly expensive.
The latest Nvidia L4 GPUs (24GB) instances are currently less than 15c p/h spot.
T4s are around 20c per hour spot, though they are smaller and slower.
I've been provisioning hundreds of these at a time to do large batch jobs at a fraction of the price of commercial solutions (i.e. 10-100x cheaper).
Any problem that fits in a smaller GPU and can be expressed as a batch job using spot instances can be done very cheaply on AWS.