I say this as someone who built, manages and operates datacenters and colo spaces.
I say this as someone who built, manages and operates datacenters and colo spaces.
If the problem is that your group of six-figure salary people only know how to put data into AWS, or other cloud services, and not design/engineer/maintain your own bare metal infrastructure as well, then that would definitely be a limitation.
For reference, a few petabytes of data is not actually that many systems these days, if you have something like a bunch of 72-drive supermicros or equivalent with 14-16TB drives in them. Set up properly this can be administered by one FTE (of course with additional staffing/tech resources for when that FTE is on vacation/unavailable, and appropriate training for other persons who might have admin on the setup).
my very rough calculation here says that a 36-drive ZFS RAIDZ2 composed of 16TB drives is something like 492TB (447TiB) usable storage capacity.
so five such arrays would be 2460TB.
compared to the monthly AWS bill for 2.0 to 2.5TB of data you could probably afford to entirely duplicate the whole setup in a twin identical set of hardware at a geographically diverse off-site location.
It's not about whether or not the engineers can make the colocated setup work.
It's that you're going to pay a lot of hidden costs with a colocated setup. Engineers can't set up, maintain, and do on-call for the colocated setup without subtracting from their primary working hours.
Each additional engineer you have to hire to help with the colocated setup is $200-400K fully loaded out of your company's budget. If you have to hire 3 additional engineers to fill out your colocated on-call schedule and help set up and maintain the system, that's easily an extra $1 million per year on your budget. Cloud is expensive, but $1 million goes a long way.
It's easy to look at a potential AWS bill and a potential colocation and hardware bill and declare colocation the winner, but then you still have to set up and maintain it all as well as constantly train everyone on it.
With AWS, you can hire engineers with AWS experience and they'll understand the big picture of how to work with things on day 1. With a custom setup, you're at the whims of whichever employees set up the system because they know it best.
Colocated systems tend to work very well at first when the original engineers who set it up are all still at the company and it hasn't run long enough to start encountering rare failure modes. They quickly become a nightmare when your engineering staff turns over multiple times and nobody can remember who knows how to do what on the colocated system or if the documentation is up to date or not.
Everything above really sounds like it's just regurgitating AWS sales person talking points.
Sounds like a systemic management / CTO-level problem to me if a company isn't willing to put in place the hiring practices and compensation, documentation systems and operational procedures to deal with that sort of concern.
If your core engineering staff is turning over multiple times for arbitrary reasons you have other problems to deal with.
> Engineers can't set up, maintain, and do on-call for the colocated setup without subtracting from their primary working hours.
If a company can't hire datacenter techs to install hardware, cables, and swap hardware as smart remote hands, maintaining as little as a couple of 45RU cabinets of gear, you also have other management/systemic problems to deal with. I'm looking at this from the point of view of a facilities based bare metal ISP that owns/runs all of its own hardware, and can tell you it's not rocket science.
People leave for all sorts of reasons: Moving for family reasons, becoming stay-at-home parents, moving for a spouse's job, retiring, starting their own companies, or even just getting bored and wanting to do something different. Or it could be as simple as getting promoted to a different role.
It's unrealistic to make engineering decisions with the assumption that the same engineers will be around and stuck on the same project forever.
Like the OP said: Every hour they have to spend working on the colocation setup is an hour they aren't spending on your company's competitive advantage, so you have to hire more engineers (and more managers) to compensate.
> If a company can't hire datacenter techs...
How many techs do you think you need for reasonable on-call coverage? 3? 6? Add a manager in the mix because you need someone to manage them.
The costs add up quickly.
It's weird to see people championing colocation as a cost saver and then pivoting to arguments that you just need to hire more engineers and techs and manage them.
Employees are expensive. One of the primary benefits of cloud is that you don't have to hire and manage all of these employees to do all of these things at the colo.
I can’t remember the last time I took a call outside of office hours, and even in hours it’s very rare. There’s enough resilience built in that any issues can wait until morning
The last major outage was in 2017, before we had a third member of the team. I was on the other side of the world installing a new system, the other was on leave. We had a network issue, OSPF melted and knocked out some services, we were down for about half an hour as I rebooted the core switch pair remotely.
(We’ve since redesigned so that doesn’t happen)
We get paid nowhere near six figures either.
Sure you can be ridiculous, I remember one team I worked on that employed a full time unix contractor (on 3 times the staff wage) to look after 6 servers and deploy a tar ball every few months. I replaced him with a small shell script. Another was a DBA looking after a small oracle database (oracle - which of course is that generations “just use amazon”)
And then basic other things like having remote smart hands ready to go, and common failure items like fans, power supplies, fan trays, hard drives pre-positioned and ready to swap in. With MOPs for swapping them. Stocks of basic things like fiber patch cables, commonly used transceivers, copper patch cables, stored in every cage.
Or we have seen all of this and that's exactly why we don't want it.
Building a company is hard enough. Adding the overhead of developing, maintaining, recruiting for, and staffing our own datacenter is madness when I can click a few buttons and get the same thing from a cloud provider without hiring anyone extra to manage the datacenter.
No one is denying that a proper data center management system can exist. We all know it can exist.
The issue is that it's a huge distraction with a lot of potential pitfalls. Your network infrastructure with Cradlepoint LTE radios in the colocation cages sounds great after it works, gets set up, stays documented, and all the bugs have been ironed out. But that's a lot of hidden work that could have been allocated to launching the product faster.
If the use case is somebody developing a software product that is a totally other scenario.
“Building a company is hard enough, adding the overhead of actually building a company is madness”
> But that's a lot of hidden work that could have been allocated to launching the product faster.
The topic of this thread isn't about a brand new start-up with no resources trying to build their product as much as possible, it's about a company that spends millions in cloud invoices.
Building a company using AWS, totally legit. Scaling this usage up to millions of dollars because you don't want to manage the hassle, fair enough it's your money, I don't mind if all your profit goes to Amazon really…
One big pain that no one seems to mention are the mundane details that come with running your own DC. Even in our fairly small operation we had half of a person dedicated to maintaining the warranties on the hardware, sourcing replacement hard drives for 10yo servers, disposing of retired hardware, purchasing new hardware, etc.
These things take a huge amount of focus and don't contribute to your product at all. The "cloud" has removed all of this busywork and I wouldn't go back to the old way for almost any amount of money.
People move to the cloud to escape their company’s IT process… there might be some unicorn company out there that does infrastructure “right” but I’ve yet to work there.
Humans are incredibly expensive and notoriously unreliable when compared against “machines”; or in this case an API.
It’s usually worth paying 2-3x the cost to have someone else manage something for you with a given SLA, because that’s what it will end up being when you decide to bring it in house when you take into account the time and effort needed as well.
A “good” & “reliable” Systems Engineering team, that can offer 24/7 support will take around a year to hire and setup, and they need roughly the same amount of time to transition you off AWS in to your system. They probably need closer to 3-5yrs to give the same level of documentation, API’s, tooling, processes, UI’s and training that AWS already provides.
Let’s call it 5 years to get to the level of AWS when you started the transition.
A decent team of 5-7, including engineers + PO/PM + UX and so forth, is at least $1.5M/yr. That’s $7.5M over 5 years, not including your new hardware and networking costs. Let’s call it $10M. You’re also 5 years behind AWS now, and over that transition you’re still paying AWS, and your development speed has halved as you wait for your new team to build or transition infrastructure.
You can trade cost and quality for speed and have everything ready in 2-3 years by setting up a few teams. Add HR support, more contractors etc etc. you’re looking at a $10M+ outlay again, regardless.
Or you can keep paying AWS $5M/yr, renegotiate fees often and literally not worry about that headache and focus on your product.
> They probably need closer to 3-5yrs to give the same level of documentation, API’s, tooling, processes, UI’s and training that AWS already provides.
You don't need to build an internal AWS to manage your own servers.
For a growing company I would always weigh the organizational overhead of moving on-prem to the actual cloud costs. If your goal is in the next year to grow revenue from 50MM to 100MM, then a company wide engineering effort to move to bespoke platform to save 5-10MM just isn't a good use of time (keep in mind you are trying to hire and release new features).
It's not that it's impossible; it's that trying to both scale up your business and provide infrastructure to yourself is just unnecessarily increasing the difficulty on your businesses execution.
* https://liveramp.com/developers/blog/google-cloud-platform-g... * https://liveramp.com/developers/blog/migrating-a-big-data-en...
One big point was challenges of maintaining multiple colocation sites, with cross replication, for disaster recovery. Since Hadoop triple replicates all data within one DC, this requires 6 times the disk storage capacity of data size for dual DCs. In contrast, cloud object storage pricing includes replication within a region with very high availability such that storing once in cloud storage may be acceptable. Further, you also need double the compute, with one of the DCs always standing by should the other fail.
*) https://hadoop.apache.org/docs/r2.8.0/hadoop-project-dist/ha...
Probably some learning pains but at least i don't have to "lose" half my storage because the sizes aren't paired with anything.
It all depends on the size of the investment and how you need to run it. I built a "new" environment on a company premises due to some compliance requirements that would be cost-prohibitive in AWS or GCP. The gear was procured through a leasing vehicle, and the hardware vendor had an SLA for delivering compute and storage. HPE happened to win the bid.
There is very little difference operationally. From a costing perspective, it's about 40% less than an AWS solution. But in fairness, the customer had an existing investment in a facility - you'd reduce the savings if you had to lease appropriate space in a colo. There are some differences in terms of headcount, but those staff aren't in NYC/SFO/BOS, so they are very cheap -- senior level engineers for $80-120k, fully loaded.
Startups do stupid shit like buy supermicro computers and cobbling together hardware that gets them into trouble when the mad scientist moves on to a new gig. Makes sense when you're drowning in VC money and need to hire people, but doesn't make sense in most other ways. You avoid that by doing competitive procurements and paying marginally more for HPE/Dell/Lenovo/etc.
If you think that standards based x86-64 hardware running Linux and ZFS, or FreeBSD and ZFS is something that is super unreliable and requires a "mad scientist", then yes, you are definitely in HPE and Dell's target market.
Home built web frameworks (which apparently aren’t “bloated” and “slow”), piles of bash scripts because they never heard of Salt (or whatever is the latest config management tool)…
Almost always they think they are “saving money” by doing what they do, rarely do they ever consider the opportunity costs to rolling the entire software stack from the hardware to web stack on their own.
Good times.
> Home built web frameworks (which apparently aren’t “bloated” and “slow”)
Well, let me show you something: https://code.djangoproject.com/ticket/31624
This is a performance regression of Django's ORM for deleting rows from a table, a query that should be about as simple as it gets:
delete from tbl;
And yet it's orders of magnitude slower because it somehow generates a bloated query. So yeah, let me tell you, the frameworks that people are using off the shelf–there's a pretty good chance that they're 'bloated' and 'slow'.> piles of bash scripts because they never heard of Salt (or whatever is the latest config management tool)…
Well, that's kind of the point, isn't it? If you're trying to get a project off the ground as quickly as possible, you kinda don't care about 'whatever is the latest config management tool', you will use what is already in your toolbox that will get the job done.
btw, we also wrote a HA cluster software for sun solaris in year 2000...
And AWS systems do not have to face the same problem? Is it easier for a new staff to come in and modify a millions worth of AWS system than a colo system?
As an example, imagine if the founders and engineers of Backblaze thought like that.
* to get performant access to a storage cluster is non-trivial, there are many different variables in place which must be correctly tuned to get good performance. Network topology, high quality NICs and switches, a tuned Linux kernel, client side caching settings, network packet sizes, file system block sizes, erasure coding settings, etc.
* your solution mentions nothing of backups, offsite failovers, or disaster recovery plans.
* your solution mentions nothing of a physical datacenter: fire suppression, battery backups, hvac, power supplies, backup generators, server racks, sound isolation, workspaces for hardware maintenance, network cable routing.
* If you have multiple geolocations you need to have dark fiber or ip transit between locations with multiple ISPs to have high speed connections between sites without downtime.
In addition to the raw costs you have to factor in the lead time of building a qualified infrastructure team, building out the requirements, provisioning hardware and datacenter space, setting everything up, and then tuning everything. With infinite money this is probably still a 2 year lead time at minimum.
I do agree that running multi petabyte workloads in AWS is probably not optimal, but when you are a startup in the growth stage it is probably better worth your time throwing VC money at AWS and building out your product. Eventually, the business should probably migrate to self managed infrastructure once the right product fit has been found and the business is looking to streamline.
Last thing I remember, I was
Running for the door
I had to find the passage back
To the place I was before
"Relax, " said the night man,
"We are programmed to receive.
You can check-out any time you like,
But you can never leave! "
2PB is really not that much stuff these days. It's less than two cabinets of equipment and that amount of space (80RU or so) includes routers, switches, AC power distribution, OOB, etc.
You wanna push 1 evening a bit late to deliver something valuable for the project? Sorry, no can do.
I don't even factor in horribly expensive migration projects that brought actually 0 added business value for the type of apps we use. We still have to keep our Network, Windows and Unix admins, various App support personnel etc., there is plenty of work for them with AWS. Not 1 single IT guy was made redundant.
No cost savings, in contrary.
The public cloud and managed services work great for the most common use cases, but go outside those and you start having to engineer around limitations.
If you have a sizable footprint in any given dimension you're trading one complexity for another.
One possible solution for this if you want to do it as bare metal you control, is leased equipment (even with $1 buyout at end of term), which can be accounted for differently than purchasing it up front.
You can generally buy whatever you are renting from AWS for 1 to 3 months of an AWS bill.
The only thing I don’t get from colo is a bunch of other customers thrashing the cache on my CPUs
Databases are not web servers there’s no possible way to not run a database on smaller / fewer instances when running at non-peak times. Instant scaling is the only possible advantage AWS could bring. However with the prices they charge it’s simpler and cheaper to just buy/rent your own hardware. Especially if you have to pay egress fees. (bandwidth is really the biggest ripoff)
I don’t like lock-in, but the prevailing view has always been in favour of lock-in, be it IBM mainframes, oracle databases, windows servers etc, and if you swing that way aws has tempting offers.
Oh and databases do scale. Say you want to run end of quarter financials that require a lot of processing for a day, you bring up tons of read replicas and away you go
Also, this is generally why you run financials overnight. If your hardware is serving transactions during the day it can easily run your quarterlies at night.
nginx is far easier to maintain than AWS load balancers which is what load balancers their load balancers are. The best part about nginx configs? They are cloud agnostic and will work on everything from a Raspberry Pi to a 128 core EPYC server.
I'll tell you something about RDS reliabilty, your monthly maintenance window brings your DB down far more often than a single unreplicated server ever fails. EBS (like the entire thing) has failed more times in the last year on us-east-1 than my colo RAID.
The selling point of AWS is that if you pick AWS and it fails you can say, well the richest guy in the world can't figure this stuff out so it must be impossible, when in reality high school kids could make a more reliable system. If you pick AWS you have the unreliability of the base software / hardware of their systems plus whatever the AWS engineers fuck up. At this point it's pretty clear that they can't even keep a SAN working.
What is the amount of work required every time some team wants to spin up new stuff? What is the turnaround time between when they file the ticket and the work being completed?
You put your config in there and then you can copy them anywhere, even between AWS accounts. They’re even cloud agnostic so when you move to GCP they still work.
Files work great with fucking shell scripts, entirely compatible, it’s rumored that shell scripts beneath the hood are also files. https://github.com/brandonhilkert/fucking_shell_scripts
You can use FSS on docker, kubernetes, bare metal (otherwise known as computers), Windows, Mac, Amiga, VMS etc.
Plan9 which was made by the creators of Go works exceptionally well with files. Some people even say that AWS converts your LB config object to a config file for nginx when it’s provisioned.
If you want to get really crazy you can put all your files in one dir like /etc and then you can use this program grep to search all your configs at once. You can even use things like perl -pi to provide a programmatic interface to editing all your configs at once.
In other words your assessment would only be true if AWS had a 0% or a negative margin.