We decided to move 90% of our workload from the cloud to on-prem infrastructure
medium.com
medium.com
Nonsense. There were plenty of SaaS startups. There was even a little event called the dotcom boom all about internet companies. This lack of history and experience is why new companies get into this cloud-first mess in the first place.
Cloud is primarily for flexibility in iteration, dynamic scaling, or complex configurations that would be otherwise hard to do. If you have steady-state load like this then a few servers in your colo is pretty simple and far cheaper.
Companies also vastly overestimate their scale when their entire business could probably fit on a single commodity server.
No, the most expensive part is the persons time for managing it. I can rent a monstrous Dedicated Server for $400/month from OVH, but even with a UK salary, if I have to spend more than 1 day a month on it in any shape or form (and that includes the initial setup), it's cheaper to use "the cloud" or some form of a managed service.
I used to manage multiple racks worth of servers on top of managing the 1k containers running on them, maintaining the (pre-kubernetes) orchestration software I had written to deploy containers to our servers, and still had time left over to spend the majority of my time on the architecture and project management of new projects for clients.
It also should not take anywhere near a day per server per month - if it does, then in a disaster scenario it means you're unable to recover at a reasonable pace.
It is additional time because the cloud handles the whole class of "your hardware died" problems.
I've set up more than one hybrid setup where you didn't need to know if your workload was running in AWS or Hetzner or somewhere else, so we could use AWS for elasticity and Hetzner to keep cost down.
This is not a hard problem. And if people don't have the right skills in house, it's easy to outsource this (I for one used to make my living of automating setups like this and operating them on a retainer basis).
I don't think anyone is claiming it's a particularly hard problem. It's the opposite. We're talking about 1 hour a month. Even that little effort still ~$200 a month and a $200 a month managed DB instance is pretty beefy.
Been there, done that, many times, and I know exactly what it costs people, because I billed by the hour for it, and your cost assumptions are just way off.
If this was the case, my billable hours when I was doing contracting would be about 1/10th of what they were. I lived very comfortably of troubleshooting for teams who had gotten themselves into a thorough mess with this attitude. In fact, I earned more from the teams who insisted on cloud setups because they rarely understood the operational issues with it, whereas teams who chose dedicated servers generally thought about operational concerns more.
> Moving from 1 server to 2 is an architecture problem.
Moving from 1 container to 2 is an architecture problem.
Moving from 1 server to 2 running those containers is an architecture problem with well established known solutions.
E.g. for starters you're assuming no containers. But putting the application in containers and leaving the host OS only for basic infrastructure setup is basic practice today if you're running your own servers.
Once you've set up a network boot source (tftp etc.) to network boot of an installer with a suitable setup script that you can trigger via IPMI, and a directory service and a basic orchestrator (be it Kubernetes or something else) on your network, it doesn't matter much if you have 1 server or a 100 - they come up and you put containers on them, and they look no different than a cloud service to the devs.
You're right that it's easier to take shortcuts if you have just a few servers, but it's just as easy to take shortcuts with just a few cloud instances - the number of pet containers I've seen over the years is terrifying.
> The last thing you want to find out is that someone ssh'ed in and installed a package that's required for your service to run
Which is why you don't provide ssh access to the host servers to anyone without an understanding of ops concerns, log everything, and do all updates via an automated setup of your preference, and why you regularly recycle the containers whether you run a cloud environment or dedicated servers.
The reality is that cloud systems do not at all make you immune to this - I've done year long projects to regularise AWS setups that were full of undocumented manual changes to bring everything into a terraform config for example. Often they're worse, because there are a whole lot of unobvious places to look for extra bits and pieces.
> and if you're running a big old monolith, you likely need > 1 instance for some sort of redundancy anyway.
Nobody here suggested a monolith. Nobody is suggesting you forgo redundancy. The. The point of this is that often you can put clear upper limits on the computation you will need to do for either your system as a whole or for a given subsystem, and you can guarantee that you will never need more than one server for a given part of the system.
E.g. a real example: I've worked on a system that did some processing of data about companies. We know this will always fit on a single system because the population growth of humanity is slower than the performance growth of a server and the total number of companies worldwide fits on a single system with several magnitudes to spare today, and the type of companies we were interested in is just a subset. At this point, if you architect a system like that on the basis of assuming you will need to resize, you will risk making choices that makes you far more likely to have to resize. E.g. all of the data I'm talking about can easily fit in RAM on a relatively moderate server now and forever, but the moment you start planning for partitioning the data you have added orders of magnitude of performance overhead for communication.
Properly assessing which parts of a system needs to be able to scale is at the core of architecture, and architects (or devs; it's terrifying how many places lack anyone with architecture experience) who are just planning for infinite scalability for everything is huge red flag to me. It usually means they don't understand their system. For a startup, especially, this is an existential question - preparing for unnecessary scaling has killed many startups.
A ec2 instance or other vps requires the exact same maintenance as a bare metal server. They are essentially the same except one is virtualised and the other isn’t.
The cloud actually requires more investment for large organisations. Previously you might of only had a handful of sysadmins but now you have a large dedicated platform team doing devops type work building your own abstractions/paas on top of your cloud.
The advantage the cloud has is flexibility. You don’t need to go through a lengthy process to acquire hardware. Likewise cloud services are disposable, no longer need something? Hit the delete button.
I think the more PaaS like services such as Heroku, Lambda, Fargate, Google Cloud Run do better realise the less maintenance story but not cloud generally.
Owning a BM server different toil - server parts fail, the network it attaches to needs control and it fails too, firmware needs updating more regularly than ever, DC space needs managing over time etc. For a small number of machines maybe this is NBD. For thousands of machines this is just grunt work which while automated, still needs change management and control - rebooting the whole fleet in the middle of the day definitely opens doors in your career.
Doing everything you did in the DC in the Cloud is absolutely the worst way to adopt Cloud. Owning an OS is a non-goal, you’ve gotta climb to a higher abstraction - workloads, and quit caring about machines. This is where most companies fail.
Ani. If the hardware you’re running on is dying. In ec2 you stop it and start it. It’s on new hardware. If you run bare metal you’re screwed.
Disk is dying ? Don’t matter cos your data exists multiple times over in AWS elastic storage. With bare metal you got to shut down and replace.
The cost of managing hardware is gone with ec2.
Software on the other hand yes that is the same amount of effort.
If you don't have sysadmins and network admins with the right experience, you can easily find yourself in a bad spot with single points of failure, servers that can't be easily replaced, oversubscribed PDUs, misconfigured switches/routers... and any number of other problems that aren't occurring to me right now.
We went dedicated early and while it didn't make sense at the time we now run at way lower cost than the competition could dream of.
The crossover point where cloud is more expensive is really low even if you have zero in-house experience. Exactly where it is depends on your amount of egress, as that is where AWS in particular really takes advantage of you.
Not really, this is what RAID is for
Or you can fork out the extra cost to have hot swappable.
If you're saying: "I dont want to pay $300 / month for EC2 instances at AWS when I can get the same hardware specs for $150 at X"
Chances are you're getting a shitty old desktop or barebones rack that lacks features like hot swap and raids.
Do you believe you're getting redundancy when you go with hetzner and a cheap desktop grade processor, ddr4 non ecc memory, and a consumer grade SSD?
> Do you believe you're getting redundancy when you go with hetzner and a cheap desktop grade processor, ddr4 non ecc memory, and a consumer grade SSD?
Irrespective of how much I'd skimp on the hardware, I always have a HA setup, including on EC2, so it doesn't matter in any case.
You mean, you reimage? That is the slow step, you reimage, and plug the new server. Wait a bit, and your service has one more server.
> With bare metal you got to shut down and replace.
You take the disk out and plug a new one. You don't turn things off because of a disk.
No doubt, those are costly. They are also rare (disk failure is less rare, but still rare).
No. When you /stop/ an EC2 instance, and /start/ it again, it moves. You do not need to reimage. This is even requested from AWS when they are having hardware failures and need to move customers off so they can decomission the hardware. They request you stop / start the instance, if you do not do it by the due date they do it for you.
> You take the disk out and plug a new one. You don't turn things off because of a disk.
If you have a storage array sure. But if you're getting bare metal hosting from a provider, you're not always getting hot swappable storage arrays.
> No doubt, those are costly. They are also rare (disk failure is less rare, but still rare).
It was 1 example, obviously there's many different hardware issues that can go wrong.
If you have any server-level hardware bought in the last 20 years or so, it will have the drives in hot-swappable bays. If you then choose to not set it up in RAID, it's just incompetence.
> If you have a storage array sure. But if you're getting bare metal hosting from a provider, you're not always getting hot swappable storage arrays.
If you're getting bare metal hosting from anywhere including your own colo, you have failover and the ability to order replacements while your system is still running. This is only an issue if you're architecture is fundamentally flawed, in which case you're likely to mess things up whether you're on bare metal or in a cloud.
> we tied Hetzner servers into our private cloud layer, and migrated containers and shut down servers as it fit
Building your own services on top of AWS is always going to come out more expensive. EC2 + EBS volumes alone are going to be more expensive than going with hetzner (particularly if you're not looking at reserved instances, and not utilising spot for burst). You mentioned that you are building your own private cloud layer and migrated containers; the cost of building that out in the first place is likely enormous compared to building and running on top of fargate.
At the time we didn't have a choice, as nothing like Fargate existed, but today it's also easier to do setups like the one we did. It mostly involved rsyncing base images over, rsync and a super simple storage service for backups, a LDAP based directory service, and and a thin layer over vzctl (first) and docker when that became an option, coupled with a VPN setup to tie our locations together, and a reverse proxy setup that did dynamic lookups in our private DNS fed from LDAP.
It is hard to do as a multitenant public service, it's trivially easy to do as an internal tool that needs to support only exactly what you need.
I've built out setups like this for a number of clients since, and it's typically 1-3 months of work to automate pretty much everything depending on complexity, and so it pays for itself quickly from a very low scale.
The first company I did this at wouldn't have been profitable at all if we'd relied on AWS
Completely agree here. Running an EC2 Spot instance + ebs volumes for 730 hours a month is a total waste of money, but running RDS behind fargate and ECS with an ALB is likely to save you time and money by month 2.
Definitely not true. With a dedicated server, you need to handle backup and security yourself.
Any reasonable dedicated setup will involve imaging your server, and so the OS image is not something you need to back up - if it fails you reimage. If you even store the OS image on the server at all rather than network boot.
For you to lose your OS image with a dedicated server only takes your HDD to die.
For you to lose your OS image on EC2 (where you made a snapshot of your volume in one-click) would take a lot of shitstorm to happen at AWS -- as I presume that they backup across sites.
Only if you don't have a backup.
Why in the world do you think anyone would store their only copy of an OS image on a single server?
For systems I set up, to start with, the OS is mostly immutable, booted and updated transparently to match a master image. If it gets destroyed, we just image a new server. The applications all run in containers, based on images stored on replicated file servers. If they get destroyed, we just re-deploy on a different server (in fact, automatically redeploying is trivial).
Only the application data is unique to running servers, and that needs to be backed up just as much whether those containers run in a cloud environment or locally, and again it's trivial to have automation in place for the backup and re-deployment of that. Been there, done that many times.
> For you to lose your OS image on EC2 (where you made a snapshot of your volume in one-click) would take a lot of shitstorm to happen at AWS -- as I presume that they backup across sites.
For me to lose my data on any bare metal system I've run, multiple servers in at least two different data centres operated by at different companies would need to fail at the same time. This is not hard to set up, and it's a one of setup. You then need to test your backups, just as you need to with AWS - an untested backup is not a backup.
But your assumptions of failure scenarios is also flawed. You need to protect against e.g. disgruntled employees, hackers, bugs as well. If you rely on the same security to protect your backups as your main setup, you don't have a backup.
EC2 is great when you can justify the cost, but it does not remove the need for a proper backup policy and processes to test them.
Definitely not as clear-cut as you seem to imply.
There’s more than a kWh rate for that though. There’s battery backup power, diesel backup, and failover hot switching. It’s not like just plugging in your phone to charge.
Thinking thru the factors, many seem obvious, but are often forgotten/ignored when comparing rental or IaaS costs.:
Space is a fixed cost driven by market rates and maximum Server Capital Costs i.e. floor space. Failing to fill the room increases your cost/server efficiency. Pretty common to run out of thermal/power before you run out space, as equipment efficiency increases through the lifespan of the facility.
Server and Power costs scale together, carry a minimum cost for keeping machines on, and vary based on utilization. Again if the servers aren’t doing work, your efficiency ratio will drop. Larger space typically have pre negotiated power commitments too - failing to consume that carries fiscal penalties. Servers unit costs are fairly cheap, storage not so much. Full utilization throughout capital/lease lifespan is the goal - anything less increases relative cost/server.
Network costs scale with Server Costs, and vary again by utilization - minimum invest rules apply, all servers need at least one network port, as well as upstream Core/TOR/Miniswitch gear. The network gear lifespan is typically longer than servers, but shorter than facilities. It usually incurs annual support/maintenance charges too. Bandwidth charges are variable as expected.
People costs scale with a step-function and numbers driven by minimum coverage requirements, task complexities, and level of human toil. Performing any task on a device by-hand is expensive in most markets - touches on tickets, change management, task time etc. Fully burdened S+R in Western cultures is typically ~2x the salary - a $80K employee probably costs around $150K by the time all the workplace costs, taxes, and benefits are paid. Network folk are typically premium resources compared to DC Ops. Sustainable 24x7 coverage looks like a staff of 3-4 people.
Even in the pure rental or managed BMaaS, the human cost can quickly dominate the economic model. Owning machines and OS’s is expensive at anything more than a couple of racks of machines. Eliminating people and human change/release from touching things in the datacenter is probably the first priority. Otherwise it is hard to consistently drive that human number down and meet service quality expectations for 24x7.
What is truly expensive is buying a bundled service with massive margins.
Sometimes you have the luxury of not caring, e.g. when building very high margin, low scale tools where human costs dominate due to dev or other parts of the business, but for anything that starts requiring hosting costs beyond even as low as a couple of k a month, you're leaving money on the table.
Once your past a few racks of equipment, you have generator tests and service appointments, redundant AC, Redundant UPS, dual power to each rack, etc. Dual Internet connection, and a link to your other server room that you use for DR, etc. The costs and complexity quickly escalate after a server or two.
Mind you this was only about 4U worth.
Absolutely, nevermind managing OS upgrades and needing to use configuration management to mitigate against drift. It took a lot of effort to make sure that each server was not a special snowflake that could not be reliably reproduced.
Also, dealing with vendor warranties, and being on hold with HP (or whoever), then assuring them you're running the latest firmware.. please for the love of god just replace the failed memory/disk/cpu.
I found the sheer physicality of computing infrastructure to be a source of exhaustion and burnout. It's a big part of the reason why I'm a software developer now :)
This is not always the case. For computationally-focused workloads like the OP describes, without direct customer interaction, it may be reasonable to accept risk of downtime in the event of failure. If you are doing computations that take weeks to complete, and you checkpoint regularly, does it really matter if your computation finishes on Saturday morning or Monday morning? If not, you can probably accept 6 hours of downtime once every couple of years and eliminate all of the redundancy overhead described above.
In HPC, the general rule of thumb is buy your hardware if you can be sure you'll run compute on it more than ~40% of the time.
Hum... I'd say it's much more reasonable to look at the ROI. Making investments to supply peak demand or to hedge against rare risks is perfectly ok.
I started an ISP in 1996 and had the same reaction you did to that statement in the article.
The only thing I’d add to what you said is the article said what drew them to the cloud - hefty free credits.
Still sometimes reasons to use colos, but I think it's important people consider that the choice isn't just cloud or your own equipment on prem or in a colo - managed servers can get you most or all of the savings too.
Agree about scale. Most software devs have no idea what can fit on a single server, and sometimes tend to start wanting complex scaling solutions for things where every possible customer they could ever get could fit on a single server
It took longer to provision than EC2 and all you had for storage is a fixed amount of disk, but that was absolutely sufficient for many businesses.
Getting hard assets is a core part of business in basically every industry, so I found it strange to claim that it was some major obstacle just because it happened to be servers instead of trucks or factory equipment.
Of course since then things have improved and you can now expect your sever to be provisioned within minutes, along with a nice dashboard to manage it. It's a bit more involved if you want to build your own server and put it in colocation somewhere, mostly because that involves being physically present.
Edit: Actually I remember using actual cloud provider back in 2008 - it was called Virtualmaster, one of the first cloud providers in central EU. They offered a free (!) VPS with 256 MB RAM and 1/4 CPU - I had a minimalist Debian image that allowed me to run full LAMP stack on it and an IRC client in tmux session.
Another cool central EU provider was 4smart. Their offering used to be that you only paid for resources you actually used - but for VMs. If you consumed only 10MB RAM, you only paid for that. I had servers cheaper than 1 EUR/month running continuosly with their own IPv4 there. They changed the pricing structure after few years.
yes, but I am glad about how many job titles are obsoleted for many businesses due to how compute instances are managed.
much smaller organizations used to need a full blown database administrator or two, and other personnel dedicated to keeping the server up. or you were doing it all yourself and spending your time on that.
much higher barrier of entry than today where you have an untold number of computers spun up for you in an instant and a bunch of cached versions on yet more computers in the CDN and another process keeping those caches updated, with you just thinking its one single instance used because you're on the hobby plan.
When I worked at smaller places, developers handled infrastructure and development. As the company grew, dedicated specialists came onboard to help.
Today, at smaller places, developers handle the cloud infrastructure, and as the company grows, they bring on dedicated specialists to help.
The biggest difference, I think, is that we have so many specialized products. We are no longer trying to figure out how to make a shoehorned relational DB scale, instead we start with a database designed for specific workloads.
Emphasis on backup.
I disagree with this statement. Yes, you could rent, but not by the hour and based on compute power, and couldn't rent extra storage again by the hour and by the GB. Plus, you couldn't interact with these "virtual servers" through APIs.
I was at AWS 2008-2014 (early days!), and I think you should consider the impact of the "on-demand", API-by-default, nature of the AWS offering. Oh, and don't forget that with a valid credit card you could be up and running in literally minutes, not weeks.
Back then AWS had decent performance, but it was pretty bad when compared to more traditional Colo offerings; but in regards to the above aspects it dominated the scene, undisputed. IMHO, that's what gave AWS most of the initial traction.
Your disagreement here is without merit. You are talking about facets that were never even mentioned by the GP. If you wanted to list those things off as why the previous situation was suboptimal, fine, do that. But how can you disagree with a (presumably) completely factual statement?
Right - because it's so cheap you don't need to rent by the hour.
Indeed, but that's completely leaving out the single most important thing: backups. With all of the major clouds, snapshots are easy to do both at a VM level and data level (e.g. RDS), and the cloud provider takes care that the backups are sufficiently spread to be disaster tolerant.
In contrast, when you co-locate you have to take care of backups completely on your own, and with many hosters you can't even influence in which of their DCs your servers will be.
If anything, OVHs SBG fire incident should have shown everyone how hard it is to build resilient systems.
Except against the disaster of the cloud deciding it's not worth to keep you as a customer, or the cloud having a distributed failure, or the cloud getting out of business...
It's the same problem if you run your own server or on someone else's systems. You need a backup and restore plan either way.
You can make an economic argument for or against cloud in practically every IT domain, but in HPC the case for on-prem is really compelling; none of the cloud networking/resiliency value-add is relevant to batch workflows, and costs per core-hour are only remotely comparable if you use spot - which is itself a major compromise.
The only real advantage cloud has for science is object storage, which is genuinely a much better idea than trying to manage your own long-term archival storage.
If I were independent I would recommend people buy and build on-prem clusters and shuffle data out of fast scratch into Glacier, but other than that just don't worry about cloud until price pressure kicks in and we are down to 1-2 cents per core-hour on-demand.
I'd love a role where I can say these things non-anonymously, but the salary for such a position would be at least 50% lower than working for a cloud provider. Keep that in mind when talking to your supplier - we may not believe the pitch ourselves, but making it is just part of the job.
Are you saying we should stick with AWS because most stick with AWS?
My personal experience with GKE/GCP was quite good, except as expected for their support.
Broadly speaking yes - there is a lot of value in having a deeper pool of skilled people to hire from, and there are enough differences between cloud offerings to knock at least a couple of "effective years" of experience off someone who changes provider.
Ceph/Object storage comes into its own at the multi-petabyte and higher levels, which is not very many groups or institutions.
Ultimately you just have to design for what is important to you; I don't want to spend time managing this stuff any more, so keep a local NAS for my partner to access and put the bulk of my "cold" data into 2 different cloud object storage providers. Note that neither of these is actually S3; for business use I would absolutely use AWS but for personal files I can manage with the reduced capabilities and lower prices others offer.
My exposure to the actual granular costs and billing have only been limited to a small company and in that case the costs were pretty appealing compared to running everything yourself. Granted this was also a bit of a hybrid with some services local and others in the cloud.
I've not had much exposure to where the deep costs start to pile up as far as cloud services goes. I wonder where those pop up?
Outbound data.
Cloud companies generally make inbound data close to free, but outbound data incredibly expensive.
As someone who has done a fair bit of HPC I consider the real advantage to be temporary scalability. If my 'normal' compute notes have 128 GB of RAM and all of a sudden I have job that need 300 GB or RAM, with cloud I can just change a line in a config file and run that calculation on a machine with 300 GB of RAM. Or if I have a job that will optimally run on 100s of 1-core machines with only 4 GB of RAM I can set up a cluster of such machines with in minutes.
That being said I 100% agree that if you have a normal baseline workload that should absolutely be done on in house hardware.
For a lot of our customers, cloud is impossible now for HPC: shortages are so bad that you have to know someone high up at top cloud providers to get access to right-sized GPUs. (T4? Forget it -- one of our tickets is open since ~Christmas.)
We have gone hybrid, and for growing compute, going multi-cloud, with main stuff on top 3 clouds (CPU, light minimal GPU...), and GPUs elasticity on other ones. And for a lot... Yep, just buy GPUs for local dev.
I work with an academic HPC group, and because researchers generally pay only for the hardware, and maybe some recharge rate for occasional maintenance, the cost per TB per month for 100's of TB and larger systems works out to the same for Glacier Deep (about $1/TB/mo) - except there is no 180 day requirement, no egress fees, and no transfer fees. And disk just keeps getting cheaper.
I'm told that big part of the solution is their use of ZFS.
Pretty much 100% agree with this statement. Only exception is maybe at the beginning of a long series of computation, you might want to start on-demand to fully understand and size your exact needs, and then provision off-prem.
Of course, it's not globally distributed and there is no fail-safe, but for the work we're doing, that is no issue. If we employees are offline, it doesn't matter if our tools are offline, too. And now that all AI storage is local anyway, building a GPU compute node is easy. I'm still waiting for 3090 prices to drop further, though, in contrast to the article. But I also went with Ryzon 5950 and Linux. I was positively surprised that 10G fiber networking is now down to $70 for a PCIe card + 20m cable kit. My workstation now has 1010MB/s 4k random write on the network filesystem. (We used SAMBA 3.1 and CIFS mounts)
I also grabbed the python/ubuntu package lists off Google Colab and created my own Docker to imitate it, and now data processing and AI training is fast (I always get the good GPU, no luck involved) and dirt cheap. Originally the idea was to run it on OVH, but I'm now also running it locally.
https://github.com/fxtentacle/ovh-colab-sagemaker-compatibil...
Wow, that is surprising, maybe it's time I started upgrading my LAN...
10GbE baseT switches are more expensive than fiber equivalent though.
ime if your unloaded latency on copper ethernet exceeds 0.2ms per hop there is some form of powersaving involved
SPF+ switch to OM3 fiber to SPF+ PCIe card => 0.2ms
SPF+ switch to RJ-45 cable to RJ-45 PCIe card => around 3ms
Some advice: make sure you know the form factor of your NICs. I accidentally bought FlexibleLOM cards. They look suspiciously like PCIe x8, but won't quite fit. FlexibleLOM to PCIe x8 adapters are cheap though.
Why they moved to on-prem: lower and more predictable cost. At a public cloud provider, they lost thousands of dollars (of free credit they had) through "architectural blunders". And the running cost of GPU, CPU, storage, and data transfer summed up to $10K a month - at which point they figured they might as well purchase their own compute servers.
I wonder how much of this is by design. It seems it is in Amazon's interest to keep their systems and pricing as opaque as possible.
If you’re a startup that simply have a bunch of web apps and APIs where uptime and network are your major costs, on-prem is only going to become a worthless headache.
A good Systems Engineer should help to figure out such choices. Anybody need one?
Every day I deal with AWS, Google Cloud and Oracle Cloud. Previously I used DigitalOcean and OVH. I have no issues onboarding people to work with the less popular options - as long as they get how Kubernetes / Linux / containers work, it's pretty good.
EKS, AKE, GKE, LKE, DOKS, OKD, Rancher... Whatever they're all compatible with what you want to do. There are definitely upsides to the cloud, but Kubernetes is the common denominator everywhere.
Wanna run GPU workloads on-prem? Buy some servers and do so. The only hairy thing is managing your own storage, quite the responsibility. (Look at Atlassian right now).
GKE https://cloud.google.com/container-engine/docs/cluster-autos...
AWS https://github.com/kubernetes/autoscaler/blob/master/cluster...
Azure https://github.com/kubernetes/autoscaler/blob/master/cluster...
Alibaba Cloud https://github.com/kubernetes/autoscaler/blob/master/cluster...
Brightbox https://github.com/kubernetes/autoscaler/blob/master/cluster...
OpenStack Magnum https://github.com/kubernetes/autoscaler/blob/master/cluster...
DigitalOcean https://github.com/kubernetes/autoscaler/blob/master/cluster...
CloudStack https://github.com/kubernetes/autoscaler/blob/master/cluster...
Exoscale https://github.com/kubernetes/autoscaler/blob/master/cluster...
Equinix Metal https://github.com/kubernetes/autoscaler/blob/master/cluster...
OVHcloud https://github.com/kubernetes/autoscaler/blob/master/cluster...
Linode https://github.com/kubernetes/autoscaler/blob/master/cluster...
OCI https://github.com/kubernetes/autoscaler/blob/master/cluster...
Hetzner https://github.com/kubernetes/autoscaler/blob/master/cluster...
Cluster API https://github.com/kubernetes/autoscaler/blob/master/cluster...
Vultr https://github.com/kubernetes/autoscaler/blob/master/cluster...
TencentCloud https://github.com/kubernetes/autoscaler/blob/master/cluster...
These are all cloud providers that invested into their own managed Kubernetes, I'm certain all of them aren't as sleek as the big three, but it shows that there's definitely momentum behind sticking your workloads into Kubernetes.
In the past, if you needed a load-balancer & reverse proxy you'd use Nginx or HAProxy regardless of the underlying machine. Now in the cloud, although you can technically run it on a VM, it's not "best practice" and you should instead reimplement it using your cloud vendor's proprietary equivalent, whether it's AWS ELB/ALB or something else, and that experience isn't portable across competing clouds.
Those are not tools from the past. They still work really well and there is no law requiring your brand new startup to be on AWS.
Most companies (even more when they are B2B) have very predictable workloads.
But anyone who’s done infra at scale can easily get up to speed on any of the big three—-as long as you take the time to understand the differences, and actually model costs before starting to play around.
Most don't. 1mm requests per minute is very pedestrian for a single vm in virtually all cases. 1mm per second is totally reasonable too if you are careful with a few things...
I genuinely believe you could put the literal public Netflix biz experience on a single VM. Account management, billing, preferences, view history, etc. The only pieces that need cloud scale are ddos mitigation, 4k video streams and media-dense static web content. Most businesses do not have strong need for these things.
They throw one of these boxes at an ISP and interconnect. 4-6 year no touch reliability, couple hundred TB storage.
Modern hardware is quite something. They will saturate 2x100GE. That's in the thousands of concurrent streams per box.
The one weird trick that cloud providers hate.
There is a lot of elegance with this type of setup too. You can have your analytics system receive a synchronous replicated log from the production system (what else is it gonna analyze?), so it can also double as a manual failover site.
> First, it’s important to note that Enzymit’s use of cloud computing mainly entailed computationally intensive calculations for protein design. We do not have (yet) any public-facing applications that need to scale across multiple geographical zones and handle millions of requests per minute. Our primary use case is running CPU and GPU heavy analyses, and for that use case, we have found the IaaS/public cloud solution to be far from cost-effective in the long term.
A user clicking "Watch later" on a video could theoretically be communicated in something as small as 64-bit integer for the user/session id, one for the command type, and another for the identity of the actual video. With serialization, padding, etc., you are still probably well under 50 bytes for this one event.
Being sloppy with data throughout is certainly a good reason to need more pipes and servers. With enough discipline, you can process events at rates far exceeding 1 million per second with a single box and non-exotic network stack.
If you imagine that the client already has a full database of all the available videos and metadata about them, then you could get by with tiny amounts of data, but that's not even close to the actual circumstances that Netflix operates under.
I'm seeing more of these "on-prem infrastructure" posts, citing costs, efficiency, and cloud complexity. We run most of our infra on-prem, and have looked to moving a few bits and pieces to the cloud, but the math almost always works out to buying hardware. Meanwhile, I talk to some <cloud stack> friends, and their opex costs are astronomical for the traffic and size of their products.
The cloud is extremely convenient, and I would choose it if I were launching a new product, but past certain sizes and expenses I would start to do some math. It's not terribly difficult to run these cloud "shrinkwrapped" products (such as load balancers) on-prem. Things like object storage seem more difficult to me. I'm also hesitant to admin a database, they intimidate me :)
A lot of places do it and just fly under the radar, but if you're going to publish a blog post bragging about it and how much money you are saving ...
The RTX datacenter restriction as far as read, but not a lawyer, is for data center providers like aws, ovh, hetzner etc to provide servers with these gpu and rent them.
Lambda sells "GPU workstation built for Deep Learning" with RTX 3090.
Also, we serve this from a single epyc based server, using elixir/phoenix. It's at about 6gbit/s outbound traffic. I realize this is not uber redundant, but it works and keep the costs low.
We do hosing for commercial HVAC systems and due to software requirements and the human factor of training the VM's we had in azure cost us significantly more than on premises servers. Add to that we generally keep those servers a little longer than some others would just increases the savings. This is still true even though we pay a colocation to host those servers.
The cloud flexibility is completely irrelevant for us in light of multi year service contracts from each customer we work with.
I think there is still a lot of potential for open source management of core EC2/S3/networking capabilities (aka "core AWS IAAS service"). We have a fair number of cloud abstraction layers now, and obviously kubernetes, you'd think we could as an industry produce core apis for doing resource listing, availability, etc.
Maybe some of the problem is that devs have a LOT of experience with the "ask" side of IaaS: gimme storage, gimme vms, set. But they have no experience with the "provide" side, and the sort of one-off manual nature of installing networking and machines doesn't have good standardization for "reporting available resources".
At this point, aws apis are somewhat stable. (I would bitch about the error codes and documentation... but anyway). It's obviously "good enough" after 10-15 years of them.
Are there projects that try to marry an aws-ish api, which really is a reporting and request api, with a "available resources" reporting api? Are some of these things out there?
AWS ten years ago was liberating. It was progress. It was a good thing. But Amazon is not a "do no evil" corporation, much the opposite. And you see this in AWS with its treatment of startups, open source projects, and other manipulations. They are a monopoly now, or at a minimum a dangerous cartel.
A real open source alternative would be a good thing. It would be good for the rest of FAANG, it would encourage competition by allowing lesser clouds to offer core competencies that are drop-in.
The first internet boom was exactly what the author describes. I worked for an internet e-commerce start-up in 1999, and guess what? We had dedicated hosting bandwidth coming into our office, and the hardware and software in place to serve our application. Of course, one of the founders was a sysadmin, but I had friends working for similar start-ups, and they leveraged one of the many co-location providers in the city to manage all the details. Yet their employer still owned actual hardware that they could touch when necessary.
After spending $100k in a year.. After doing the math, We decided to purchase three workstations at a total cost of $17k.
One is a GPU-based workstation with two RTX 3090s and an Intel i9–12900 CPU, and another two workstations with 16 cores AMD Ryzen 5950X CPUs.
It took us a few FTE days to set those up to our satisfaction with slurm, NFS, backups, and several other services.
We noticed that our RTXs, although considered gaming cards, are comparable (if not better) in performance to Tesla V100, which some cloud providers rent at the staggering price of $3.06 an hour.
tl;dr - They did the math.And bought 3 workstations instead of 3 servers.
IMHO the headline of "Why Enzymit Decided to Build its Own On-Prem HPC Infrastructure" is a bit... stretched.
They also had FTEs with experience to configure the systems (which for a lot of us technical people is a no-brainer, but if you had data scientists with no hardware experience, that might be different)
This worked for them, and I'm happy for them.. the cloud does not solve all problems... It's simply one hardware strategy you can pick.. it's important to review the options!
If you need a pickup truck two weekends a year, rent. If you need to carry tools and stuff every day, buy!
> Warranted Product is intended for consumer end user purposes only, and is not intended for datacenter use and/or GPU cluster commercial deployments ("Enterprise Use"). Any use of Warranted Product for Enterprise Use shall void this warranty. https://www.nvidia.com/en-us/support/warranty/
You don't buy Tesla cards because you care about maximizing operations pr second, you buy them because you care about operations pr KWh.
IaaS is not really competitive in this space, I don't think... If you have access to system admins, and have a consistent work load, you could avoid the cloud, trading the cloud premium for more employees and skills in your team. This is fine, if that is what you're team needs.
The cloud is not a silver bullet that solves all companies infrastructure... But they have a very profitable space, especially in small businesses, or businesses that benefit from multiple data centers. AWS simply made it easy to scale up and down, as well as scale around the world... if you don't need to dynamically scale or easy access to multiple data centers, the cloud begins to lose it's best (cost-effective) competitive edge to self hosting... Though at the micro scale, the cloud can do dead simple basics for free -- or near enough (e.g. static websites), which is fun for personal projects
Clouds also require sysadmins, they're just called "DevOps engineers" now. Those YAML & Terraform files aren't going to write themselves.
People who have never managed servers imagine them blowing up once a week or something.
Is that easier on-prem? I was under the impression it was even more difficult--especially with shared tenancy (how much incremental cost does App A auth add to our Active Directory deployment?)
You're going to need to know server power utilization under load to calculate power/cooling costs and probably some additional data on network utilization to figure out incremental costs for that
Maybe if it's a colo or managed data center that gets rolled up for you, but if you're managing yourself, you still have to figure it out
Blog post also doesn't mention cost of downtime (maybe not an issue for them) or a metrics solution (you usually get basic machine and service metrics for free on the big cloud providers)
2) Onprem experience looks bad on your resume, compared to cloud experience
3) Incentives work for people, and resume-driven development is key, as our industry is very stingy in passing the savings from 1) to the developers
So they got a very specific workload that is easy to run on any infra including on-prem.
Or not even on-prem, just renting some physical boxes from some place where they're only going to have a basic markup.
You have to wait a week to get few more boxes but that's not a big deal.
This is the kind of thing I'd imagine a corp would start doing pretty early, because the 'flexibility' of IaaS just isn't worth the cost.
> "...off-load many of its processes to an on-premise private cloud?"
On-Premise Private Cloud
I dont know why reading this my brain just cant compute. You mean you have three consumer grade computer with Dedicated GPU running 24x7?
10k of infra costs are not a lot of money in a business context.
A person to operate a private cloud with OnCall, backup hardware, capacity planing etc. costs what?
I would always try to have GPU on prem as those prices are quite high bit others I would use managed.
Cloud providers are just much better in operating infrastructure ad normal ops.
Why not share the math? Isn't this article supposed to answer why it's more economical to use on-prem?
Well, for one, because it's an exceptionally rare English word that is singular in meaning but plural in its usage. In addition, there is an unrelated English word that is singular in both meaning and construction. That confusion would naturally arise seems almost pre-ordained.