Did the math and assuming 100% util and equal performance (which is certainly not the case) payback on your Mac is 9 months...
4x NVIDIA A100 at lamda labs is $4.40 an hour and I really have not had an issue getting them.
https://www.reddit.com/r/LocalLLaMA/comments/16o4ka8/running...
OP implied that there were workloads where it out competes renting in terms of cost. Was hoping it was true for something than a single user interactive session (which can be done a lot cheaper)
When running an ML workload, the Nvidea A100 has massive GPU compute resources, and a large amount of GPU local high bandwidth memory, so it's ideal, but is nowhere near low cost.
A consumer Ryzen chip is inexpensive, but lacks in both memory bandwidth and GPU resources.
The M2 Ultra has access to way more RAM than consumer GPUs, many times the memory bandwidth of a Ryzen (800GB/s vs Ryzen 7 1800X at 40 GB/s) with a large amount of local GPU resources.
Even stepping up to a Threadripper Pro would only get you a quarter of the memory bandwidth, and those aren't exactly cheap either.
The Apple Silicon is great at really low power work, but if you dial desktop or server GPU power limits down they also become quite efficient. The marginal cost of electricity is cheaper than buying more hardware, so nVidia and others run their parts deep into the diminishing returns part of the curve to maximize performance at the expense of power efficiency.
A top-of-the-line Mac Studio will give you 192 GB of unified RAM in less than $7000. Meanwhile a H100 from NVIDIA with 80 GB of VRAM will cost you like $30000...
192GiB RAM is enough to train or inference Falcon 180B in RAM at 8-bit resolution.
The wisdom is that cloud providers are better at infra than you, and that the economies of scale make it better to piggy back on what they’re doing, but… AWS is the most profitable part of Amazon for a reason. They’re overcharging you.
When you look at the cost of the hardware + hosting. Yes, it certainly looks and feels that way.
But if you've dealt with corporate IT, and had to deal with 3-6 month lead times on getting hardware, or politics to get your hands on hardware to get stuff done.
AWS is cheap. It gives you velocity.
If your company is large enough that it can offer the elasticity of resources that Amazon offers or even 1/4 of it... and you have an IT org that will let it happen. Yes, AWS is a waste.
But with AWS... when a project dies, you can wipe its costs out, people won't hold onto hardware so they have hardware for the next project, etc...
Trust me. I've been IT, I can spec and build rack systems. I am a software dev. And I've been a dev most all my career.
For 90%+ of orgs... they don't have the maturity and skills to handle that type of infra without substantially distracting from their primary business.
Also, maintaining servers is not hard at a proper data center. It is often more hands off than the migrations cloud providers force on their customers.
It’s not like process disappears just cause you’re not on your own hardware. Infra is still its own team with its own budgets poking and prodding at every damn turn for every little thing till rejecting your requests, you escalate and then have a 4 week battle over needing the space.
If you are paying the price of being on prem, which is really lack of ability to provision and de-provision infra quickly. There's little point to the cloud, unless you just have no infra to begin with (small companies).
I'm in a small firm now. I can't imagine having an approval process to spin up a few instances to run my tests and spin them down after. That'd be silly.
Most firms don't. Or don't have the skills.
Also, the cloud can help an IT project recover from errors. Let's say, I'm about to buy 500k of hardware to setup some storage. I get my requirements, I architect it, do my design work, and then buy the hardware. I have to over provision a bit because of reality and human error... But when I discover that the requirements, shift 2mo in my project, and I've already ordered the hardware... I may be hosed.
This isn't hypothetical, this is what happens. Things evolve and shift. The cloud allows for more agility. If your firm is large enough, or has its stuff together enough, go for it on-prem.
I've got 20+ years on prem.. I've seen it fail all over. I've seen cloud be a mess too. But if you told me to clean up one. I'll take the cloud.
It’s still a win for a lot of use cases and I still do it quite often, but the meme that it’s this “click and you’ve just hired the best ops team in the world to work for you” and so the 50-500% markup is actually a bargain is horseshit. A Bizon box in your living room fucks AWS up on flops/$ on most instance types and pays for itself in 30 days.
It is one of the best ops teams on Earth: but they’re working for you like the Google search team is working for the user.
Your application isn't going to magically become HA/DR. You still have to make it that way, from your application design/coding up through the deployment.
I mean, if you're not storing your session IDs in a data store that's reachable by all the nodes behind your load balancer then no amount of infrastructure is going to save you.
That’s a realistic scenario no matter whether you’re bare metal, building out your own cloud, or using someone else’s. No amount of AWS/GCP/Azure/et al marketing changes that.
Yes, you have to learn things to goto the cloud, and I won't say it is all roses, it ain't. But... AWS is less likely to fsck it up.
If you have the constant load to burn the flops 24x7x365... go for it. If you have the ops team to do it... go for it.
If you don't... take a bit of time and learn the cloud which is much easier than getting on-prem right.
Especially for smaller firms, this isn't even a close call IMHO.
I’m quite capable of setting up whole server stacks. I did it for years, but I stopped, some time ago, and consider myself to be, for want of a better word, incompetent at being a modern admin.
I think I’d screw the pooch, so I prefer that someone who does it every day, handle it.
But I write Swift code, every day, so I’m not incompetent at everything.
If I wanted to host a website, sure, I can build a server out of parts and negotiate with my ISP and get a business pipe and handle all caching and such. Or like I can pay a provider $5/mo and get better performance and reliability with no management overhead. Yeah, maybe over 5 years I'd save more money doing it myself... but it's not worth the time.
If I wanted to generate a photo or a dozen, or a few paragraphs of text, that's like a few cents worth of cloud AI. Maybe low single-digit dollars. Or I could spend thousands on fat GPUs or a Macbook, spend forever training it, and still end up with a sub-par result.
AWS is profitable not just because they're overcharging you but because they are providing a hugely useful service for millions of businesses that don't want to deal with that infrastructure themselves, any more than they'd want to manage their own plumbing or electrical grid or roads and bridges leading to their office. DIY makes sense if you're doing it as a hobby or if your scale is so big that you would incur significant savings to in-house it, but for millions of small and medium businesses, it's just not the most practical approach. Nothing wrong with that.
I mean, it's like saying development is such a lost art... why hire a dev if you can learn to code yourself? Sure, but not everyone wants to, can, or has time.
My employer has generated its own electricity and steam for decades.
For a small business - different story.
Still haven't started on the silverware or shoes yet.
I do agree with you though. If you are a non-tech company or a company that lacks the human resources you might as well go with the cloud.
Are we in the business of building, maintaining and operating <thing to build> or do we want to buy that as a service instead and focus on our actual core business?
There's more to the cost of building and operating than just the hard costs.
Retaining good modern IT talent is getting harder and harder - and I'm not even talking about salaries.. You need a whole department including strong leaders who can hire, train, and lead the right people, etc..
This is something most companies wouldn't even know where to start with.
But if you can motivate that same sysadmin to spend his skills on something more directly benefiting your company, then you should still buy it in.
Given how popular homelabs are, I don't think this would be too hard to find.
You can throw a bunch of boxes in a closet and it'll work. A surprisingly large amount of the early Internet was "a spare box under my desk."
The problems start when they become part of your critical path and you're on vacation and nobody knows WTF is happening.
I mean, it's a risk. If you're OK with that risk then go for it.
It's really about the politics of your office.
If everyone is OK with the idea that the box is in some closet somewhere that's fine. I've been part of a bunch of startups where we were running infrastructure on spare hardware. Sure it's not HA, but we didn't need it...or it was at least HA enough for what we needed.
And I think in most cases companies want to focus their employees and efforts on their core business, and if that doesn't include setting up and maintaining hardware in the long-term, then you don't build, you buy.
If your operations halt until some poor sysadmin has to drive to the colo, you are absolutely doing it wrong.
I remember twice switching from an in-house jenkins/teamcity/whatever type of CI to Azure devops and the thing I remember the most was how much longer it took a build to complete as well as the massively longer time downloading a build from Azure vs from within the office. Even when working from home the on-prem stuff was faster.
The thing is, the build/devops teams seem to be about the same size in both cases. It's just kind of worse in pretty much every case when we do CI in the cloud.
Notes
- My experiences are largely for game development so the build times and artifact sizes can be quite large.
- I've only ever had CI/CD experience with Azure, I've not tried other cloud providers
- Since this is game development and we're using CI downtime is more acceptable than other cases. That said, I don't remember much downtime when I was working as a build engineer. I have seen periods of 1-2 hours of downtime once in a blue moon but then again I've seen that with Azure. In both cases it wasn't so much the setup but a build script deployment issue.
Also being able to cool off in the rack room when it's a hot day is always a treat :)
Meanwhile a Mini cluster is literally a bunch of mini pcs in a rack, and idk if Apple even supports this kind of industrial use. While it's a quality product the Mini isn't really designed for the datacenter.
I think they know of it and tacitly approve of this use case, as evidenced by the Mac Mini having the same form factor for ages. They’re well aware that a lot of people use Minis (and Studios now) in data centers, and that the Mini footprint is sort of “standardized” at this point.
https://en.wikipedia.org/wiki/Xserve#Intel_Xserve
But since Apple discontinued Xserve and macOS Server, they seems like don't care about this business anymore.
Mac OS X Server was its own operating system originally. It was still the same core OS, but had a ton of additional servers built in. Non-exhaustively, they included IPSec VPN, email, calendaring, wiki, SMB and AFS file shares (including support to act as a Time Machine backup destination), LDAP, DNS, and software update caching before it came to macOS proper. The Server app released via the App Store was a shadow of Mac OS X Server.
These were quite popular in small professional offices like law firms.
(Not sure what differentiates the later model Mac Mini Servers from the regular Mac Minis, since Mac OS X Server just became a $19 App Store purchase, and optical drives were no longer a thing in Mac Minis)
They discontinued the Mac mini Server line in October 2014, which was still sold with two drives instead of one. Configurable to order with SSDs by that time.
Got three databases up and running too. It's a beast. I'd definitely consider self-hosting with a few Mac Minis, that would be fun and they're really cute, sleek devices too. I paid $650 for it and consider it a great deal. Definitely should've gotten it with more than 8gb of ram but I got it to try it out and haven't yet really needed to upgrade to a unit with more memory.
You can run them at half the power usage and only lose a fraction of the performance - at least in gaming. Try for AI tasks.