The cost of cloud, a trillion dollar paradox
a16z.com
a16z.com
Seems similar to office space - a 5 person company will find the economics of coworking space compelling, a 500 person company will do better with a long-term lease with a commercial landlord, and a 50,000 person company might have their own property management team in-house. That doesn't create a paradox in the commercial real estate market, just different solutions for different needs.
Why would this happen? Because what the company builds is a large and growing money-making contraption right inside the coworking space, also using parts of the building for critical functions, and the door is intentionally kept small. The only way to move that contraption is to dismantle it and reassamble in a much cheaper, purpose-built hangar, rebuilding some of the critical parts along the way. During all that time, the contraption would stop making money.
That's the beauty of AWS business model: it's a no-brainer for a startup to use it, but the startup grows, more financially efficient infrastructure options become unattainable because of the very high cost of the move. This is the best-executed vendor lock-in I know.
Some commerce did need to scale fast and easily, and AWS provided that with great strength. But I will add that the management and executive branch were constantly and relentlessly seeking to undermine the bargaining power, day-to-day operations authority, and pay of skilled engineers who could do that kind of setup. Not that it was easy or risk-free .. I did see some powerfully strained eyes and nerves on those individuals who built and ran those kind of servers at various times. But do not leave out the manipulative and gaming motives of executive management to dumb-down admin requirements, seize authority via legal means, and seek the lowest wages across borders while doing so ..
At which point, as you say, the cost and complexity of a move is huge so you are stuck with a recurring bill.
Another interesting thing is that you pay for 100% of a VM virtual CPU even if it is only utilised 40%, but the cloud is selling your other 60% capacity to another user.
They have no incentive to charge you “based on CPU usage” for VM’s that are always on, even though it is possible. I think this would be a better approach than Lambda - to dynamically add more cores to your OS as you need them.
AWS don't overprovision CPUs outside of the t instance range, in which you get vCPU credits, not full vCPUs. If you rent an apartment but don't use it, is it the landlord's fault you don't? The landlord can't rent the same appartement to another client.
Especially on AWS/GCP/Azure, where there are many alternatives ( Lambda, Google Cloud Run, Google App Engine, etc.) where you pay per compute used.
One hardware core can be running 5 VMs at 20% capacity, and each VM owner will pay for 100% of their CPU capacity right? So in this case the landlord is renting your unused apartment to other customers then evicting them faster than you can open the door.
My point is cloud is benefiting from that ability and not passing it on the customers, which they could.
Its fully wrong for AWS, where a vCPU you rent == physical core/thread.
Are you certain of this statement?
[0] https://techcrunch.com/2017/09/15/why-dropbox-decided-to-dro...
tl;dr: to save money (and to fix issues and irritations with AWS)
Another good one is microsoft initially tying excel, word, access, etc to windows. And then IIS, SQL Server, etc to NT. An entire generation of businesses and managers got tied to windows this way.
In a way facebook and social media is a form of "vendor lock-in". The network/community itself and your "investment/time" in facebook makes it costly for many to move to another "facebook" or give up facebook altogether. The more time you spend on facebook and the more people you bring onto facebook, the more "locked-in" you are. It's quite ingenious and I suspect facebook understood this very early on.
But you see part of the magic trick was Amazon spent years having companies like Garnter proclaim that cloud was cheaper, and nobody could possibly do on-premises IT cheaper because of economies of scale. As a result you've got CIOs everywhere trying to make a name for themselves by driving their company to the cloud to show immense cost savings. They don't have time to be bothered by the actual financial models or real costs. By the time finance realizes what happened they'll be on to the next gig.
It was honestly brilliant, the number of "cloud first" strategies that originated in board rooms filled with people that don't know the first thing about IT is kind of disgusting.
Small businesses don't care because they are too busy surviving.
Medium businesses don't care because why break what works?
Large businesses don't care because their managerial careers promote short-term priorities and marquee "transformation" projects which are buzzword-driven, and cloud is the current one.
Life goes on. It happens all the time. The trick is to learn and be smarter next time.
There are off the shelf things for, say, S3 and ALB which are entirely workable, but once you start getting into more complicated stuff (like S3's new consistency semantics, or SQS) then you're looking at a whole small company (at a minimum) worth of additional work. It's a non-trivial expansion, even for a large org with lots of money.
You can avoid using these sorts of services to maintain flexibility/independence, but you lose out on their unique benefits. This isn't an accident. It's not like you're going to be able to selfhost an SES clone and get the same kind of deliverability percentages as AWS netblocks, no matter how many engineers you throw at the problem.
Some of the goods you just can't get anywhere else, and they know it.
Aurora - tidb or vitess or cockroachdb or citus etc etc
Glacier - there is backblaze, also onprem vendors
Sqs - just use Nats
Having said that, and please keep in mind that I do not have an idea much about hardware so this is a guesswork; more of a question than a statement..., I do think that the Graviton price vs performance might not have a direct benefit outside of AWS. It's cool for AWS because they have less stuff to manage at that scale and it might be cheaper for them to buy and iterate because only they use it. It looks great to the clients because the perception is that the value is better. But AWS runs 80 AZs in 25 regions so it benefits them at that scale.
Realistically, even if one had couple of hundreds of racks, does it matter at that scale?
SQS is a message queue service. It's a bit weird to claim that as "simply aren't directly available as selfhosted replacements", a message queue is pretty basic stuff.
I'm super into self-hosted shared-nothing queue services, but what AWS is doing with SQS is anything but basic stuff.
It's not limited to some. You absolutely need a whole new team (or teams) to handle your infrastructure, your high availability needs, and also your security.
The nifty serverless offerings and other features are just nice-to-haves in comparison to the core infrastructure work that cloud providers put into your system to keep it running.
Just because cloud providers like GCP and AWS and Azure and etc have everything put together to let you setup your whole infra by running a small script that does not mean nothing needs to be done in order for that to work.
My company recently shifted from on prem to cloud. And we aren’t primarily a software company (although, much like most companies, we are increasingly turning into one).
We outsourced our data center needs, and I don’t remember us ever facing these middle of the night massive breakdowns (we had load balancing between a couple of data centers) and the number of ops issues hasn’t reduced a bit since the switchover.
I do think it’s been easier to improve performance for non US customers, because spinning up servers in data centers across the world seems a little easier around the world, but the millions saved by self hosting would have easily paid for someone to spin us up a Singapore and UK colo servers.
Consolidate them, and a knife fight happens between middle management machiavellis and the one that wins is not necessarily the best on price or features. It's the one who is best at winning reporting hierarchy wars.
Instead they should pit these groups against each other as internal options for hosting to various services that provide what is needed for growth, adaptation, or cost.
AWS should simply be one of those options. Really it can be the prototyping phase, and then budget squeezing will move the app to other (hopefully cheaper) inhouse options.
What is frustrating from my years in the big corps is that AWS was cheaper than the internal "clouds" or other relationships.
You can very much do that while staying in the cloud. It's all about engineering for flexibility.
Even small companies have multiple engineering teams, and each of those teams has particular needs, and each of the engineers, analysts or general users may rely on some service AWS, GCP or some other cloud provider has. Any single person can now break this ideal flexibility requirement if it makes their life easier.
> It's already demonstrated that the big cloud providers will cut you off with little warning if you run afoul of their political sensibilities. So the only way to use those services is in the most generic way you can, so that migrations don't cost more because of it.
You are stating that you moved from one cloud to another. The question is how you manage to build a company of any size that can just flip a switch and move to a new service provider without needing to transition teams to using new services and that costing money. If we're talking about moving a simple server from AWS to GCP, sure that's no problem. But if you or anyone at your company is actually using a cloud providers services for more than hosting a server, there is no possible way you can just flip a switch and migrate over.
You’re not getting retail price.
So you’re still spending more trying to maintain “cloud agnosticism” by hosting everything on VMs instead of using managed services.
Like Jobs said, Dropbox is a feature not a product.
For the same price you pay for Dropbox and 2TB, you can get the full Office 365 suite plus 6TB of storage or the entire GSuite.
And do you have multiple redundant DCs? Are all of your database writes duplicated 6x across three DCs like Aurora?
That’s just silly - dropbox is profitable if you exclude stock comp unlike box inc (who afaik runs in cloud) and also has higher margins (big shocker here). From techincal side i hear it worked out really well for them too.
> Like Jobs said, Dropbox is a feature not a product.
What does that have to do with anything? Not helping your argument at all
> For the same price you pay for Dropbox and 2TB, you can get the full Office 365 suite plus 6TB of storage or the entire GSuite.
Should be painfully obvious to anyone arguing in good faith why that is. Also do you by any chance think google runs gsuite on aws or gcp?
> And do you have multiple redundant DCs? Are all of your database writes duplicated 6x across three DCs like Aurora?
Yes, you run mutiple redunant colos/dcs and peer them. You’re making it sound like rocket science but it really isn’t. If Internet Archive can do it with their resources (not a dig towards their talent at all) so can unicorns flush with vc cash who supposedly hire “the best”
How many people who are comparing their costs to using a cloud provider are ignoring redundancy?
And I’m sure your average startup has the capability to maintain their own redundant data center. Besides, does it add business value? Would they really be better off hiring multiple infrastructure people or just paying a cloud provider.
It doesn’t matter why it’s the case. People keep bringing up Dropbox as a shining example of moving away from a cloud provider - and still not being profitable. But ignore a little company - Netflix - that did just the opposite.
Do you not use IAM? Cloud specific infrastructure as code (again TF has cloud specific provisioners), permissions, compliance, training.
I've had this discussion many times where I work, and we've refrained from using cloud services since they are extraordinary expensive compared to what we can do internally.
Yet, people seem to be convinced that prices are ok, and that going outside will let them reduce their infrastructure footprint.
For some reason I don't quite understand, people tend to be attracted by the cloud, even when it makes no sense economically whatsoever, and talking sense into them is surprisingly difficult.
So yes, it seems obvious that people will choose the most cost effective solution. However, it seems like they don't!
In my experience when a large corporation moves to the cloud it has less to do with pricing and more to do flexibility. It's easier to get a budget from IT and do whatever you want in cloud, then to have to wait weeks/months putting in a PO for hardware, getting access to those machines and installing what you want.
And if you're wrong about those estimates, the costs are real. Overinvested in hardware upfront and your project will never break even. Underinvest, and when the marketing campaign kicks in and you get flattened on cybermonday, there is no limit to how much revenue you might miss out on.
I'm willing to pay more money to reduce the downside of my being wrong. Cloud services offer so much flexibility that I don't even have to make guesses about some things beyond my information horizon.
On prem might be cheaper but it is far from an equivalent service.
Devs look at Digital Ocean or Hetzner VM costing $10 for a ton of RAM, storage and bandwidth when the same thing internally in a bank or other big enterprise can cost $100-150. AND be delivered in 3 months, if it's ever delivered.
That's where you get it all wrong. It is not about the price, at all. It is all about avoiding large upfront capex with a small periodical opex, and in the process have virtually boundless growth potential.
Think about it: when you happen to need a bit more computational resources to run an app, is it easier to convince your boss the need to pay, say, 100€/month for extra VMs, or is it easier to convince your boss to shelve $2k to buy an extra rack? And how many times are you willing to have that same conversation with your boss whenever you run out of computational resources?
100% agree on shifting capex to opex though, capital works teams are super in favour of that, oddly enough.
No, not really. You can hardly find old used servers for less than 1000$, and you need to use beyond a a1.metal instance with 16vCPU and 32GB of RAM to even come close to that cost
> Cloud stuff only really makes sense when your resource requirements as super spiky.
No, not really. Cloud stuff makes all the sense in the world if all you need is a VM connected to the internet working 24h/day without problems. You can get that for less than 5$ a month, which would require 3 or 4 years to pay back if instead you waste your money on hardware.
> You can hardly find old used servers for less than 1000$
I picked up a second hand Dell PowerEdge 610 with 96GB of RAM for $400 AU, it's overkill by a large factor for what I need but damn. Either I was super lucky or servers aren't (weren't?) expensive.
> You can get that for less than 5$ a month pocket change and host it
Until you actually use it and then it ramps up rapidly to the point where you'd be far better off just buying a server.
And good luck selling to enterprises when the vulnerability scanner lights up like a Christmas tree with critical and un-patchable flaws on your switches/routers/UPS/PDUs/storage arrays/servers/door locks/cooling system/fire suppression/cameras. Double for the DR site. Because you need all that stuff, and you have to maintain it all 24x7 which means lots of salaries.
Yes, it's not a paradox at all. Use this thing until it doesn't make sense, and then don't use it anymore (or be comfortable with the lack of optimization). This is not a paradox.
All this happens willingly unless growth happened with your eyes closed, or really had no choice keeping up with it. Sure, use the cloud, but don't use every cloud-proprietary convenience. Use opensource software in a near-standard way. Even if using the cloud-vendor's offering, refrain from using proprietary extensions. If your plan is to grow and get out before things stabilize, then that's on who's come onboard, not resisted the conveniences, and remaining.
Moving office is relatively simple, you call a moving company to lift all your chairs, print new business cards, hire a bigger maintenance crew and it's done within a week.
Moving cloud could in worst case mean a 6 month complete rewrite of your core product depending on how much dependent you are to CosmosDB, Lambda or other shiny AWS stuff.
You could say buying a machine you previously rented is the same thing but when looking at code migrations, rewrites and launching new infrastructure cost time from R&D, not just money.
Most of the servers, each of which handles an area of the virtual world, run all the time, regardless of user load. This is very different than web serving. It's the worst cost case for AWS, where the real benefits come from dynamically provisioning a changing load. AWS, used for dedicated servers, can cost 2x that of owning your own. One data center operator says 4x.[1]
Before the conversion, I'd asked the responsible VP if they were sure AWS wasn't going to be more expensive. He retired when the conversion was complete. Now we hear rumors of the company finding out that AWS costs are higher than the old data center.
[1] http://www.happi.io/aws-instances-versus-physical-servers/
Another example is EA hosting all their Battlefield servers on AWS. Can't imagine what this is costing them for subpar performance (rumor is they had to deactivate their anti-cheat system fair fight to get acceptable performance).
Three reasonably good free-to-play indy games were built on it. They all went broke within months, because the operating costs were so high. Now, as far as their website says, they have no real customers. But they have enough money left to fund three in-house game studios, in hopes of getting a hit.
They did license the technology to the NetEase unit of Tencent in China, where it was used for Nostos, a MMO with great art that a good player could finish in a few hours. It began shutdown on March 17th, 2021.
I'm interested in seeing the Metaverse of Ready Player One/Snow Crash happen. Second Life is the high water mark of that idea. Most of what's come out since has been worse. Facebook Horizon, anyone? Virtual worlds turn out to be expensive to run from a compute standpoint. Most games are able to offload much of the work to the clients, plus much of the asset processing is done once when the game world is built. Fully modifiable multiuser virtual worlds have to do much more server-side, which costs.
On the other hand, the Well, mentioned on HN today, once cost US$6/hour. For a text only social network. Compute progress marches on.
Incidentally, Linden Lab is still looking for a VP of engineering for Second Life.[1] As a user, builder, and someone coding a third party client for it, but definitely not an employee, I'd like to see them get someone good.
I might even wager that such a comment will appear as a sibling to this one within the next 30 minutes.
They likely also retired, per the parent comment.
Everybody loves to cite DropBox here. The greater arc of DropBox is their product stalled, they lost the enterprise of their market to Box, and they found themselves a commodity in a commodity market. Heck, maybe they'll go back to Amazon if they keep doing things like this: https://aws.amazon.com/solutions/case-studies/dropbox-s3/
But that's not all. It takes a top-level directive to repatriate an entire SaaS. That's at the cost of other top-level projects. It's wild to me that any company that has significant fuel left in the tank would buy back single-digit COGS percentages instead of investing in product that could add double-digit growth for several years at scale.
Most businesses who are operating in the cloud are not so directly purchasing a cloud resource, layering a little value on top, and reselling it; most businesses make their money selling something else, and using cloud compute/storage/bandwidth/infrastructure is not a transactional cost.
For every gigabyte of storage space dropbox sells, they have to buy a gigabyte of storage from someone. And if you're looking ahead to the future of selling lots more gigabytes, you're quite motivated to find a cheaper way to buy them.
So I am never convinced by the argument that because Dropbox left the cloud, that means other businesses should.
So not sure if it's so much that the product stalled, but that leadership actively didn't want to move in that direction for too long, then lost the advantage.
Instead they spent a bunch of energy on Carousel and Paper. (I don't have any idea how much of their eng team was working on those things)
Dropbox is worth about 3x Box as of today
Its good that we have people focused on lowering cost and optimisation, thats why people in India can afford a smartphone.
On the other hand, hisotry is filled with features and products that never found their niche, from random dropbox features to thousands of products google created and killed
Yeah but why do they have 34 PB of analytics data?
The most powerful tools we have are aggregated reports and data retention periods, the same since the 1970s.
Running 10 copies of the same report (but each one with slightly different bugs) across 34 PB of data for 10 different departments is absurd, yet here we are.
Specifically for Dropbox, they're not a retailer like Safeway with 10,000 products and 100's of locations to crunch product and shipping data for - Dropbox could use a single server for their entire digital products data warehouse.
(I operated a data warehouse myself (literally) at Yahoo - 10 servers generating canned web reports for 500 managers. I heard they replaced it with a 1,000 node Hadoop cluster later.)
There is a learned helplessness when it comes to companies running their own data centers that has become widespread over the last decade. What used to be a fairly mechanical process has almost been mythologized as some kind of arcane art beyond the technical ability of any company that isn't Google or Amazon. Designing data center builds isn't difficult, it is a pretty straightforward albeit detail-oriented blue-collar engineering skill, but it seems few people learn it anymore.
Your pitch emphasizes features that make servers easier to manage, presumably lowering (but not eliminating) the cost of personnel, so you're obviously aware of this, but I'm not seeing any "net of additional, fully-loaded personnel costs" in their analysis.
It should still be cheaper above a certain scale, but the breakeven where it will make sense to "repatriate infrastructure" is much higher if you include hiring people, and will go even higher if the cloud providers run the numbers and match most of the savings with scale-based price breaks.
My team had 2 guys doing SiteOps. They would travel to the various DCs in the Bay Area and Virginia and do all the maintenance, new installs, etc. And sometimes we'd lean on the colocation remote hands to do a few things.
We had about 5 network engineers, that also handled the corporate network. (12 offices and a network backbone that connected east and west coast DCs, offices, etc).
And maybe 2 SWEs who handled things like our host OS install system, etc. Basically the next layer above the hardware.
So 9 people that were required to run all of that stuff. But really, NetEng ended up spending like 70% of their time on corporate network things because we'd add offices faster than datacenters.
So if we focused on production only, we really needed about 6 or 7 people total.
I did the math a few times (every single year) and compared our costs, including people, to the costs of moving to AWS 3 year reserve instances.
Doing it ourselves was always half the price.
Of course the difference here was we built on-prem from the start. So there was no repatriation that had to happen.
Since then, I've been responsible for large cloud infra on all 3 major providers and learned a lot about what kind of discounts you can get when you're in the double digit millions in annual spend.
I still think, at a scale of single digit thousands of servers, you'd be cheaper on-prem, fully loaded. But admittedly, I haven't run the numbers since 2017.
And those companies are in situations where either people have no idea that their teams are performing badly, or lack the political will to do something about it. Adding the services in AWS is just easier, and might be worth every penny.
When you go to a high performing firm and explain the procedures and decisions of a low performing one, people think you must be lying, but it's true: Broken IT departments are everywhere in enterprise companies, and paying double of what a good team would take to do the right thing is a bargain, given that they can't even get a whiff at a good team.
If this is the case then the CTO knows that something is very wrong.
I don't think your scenario of companies not knowing whether their teams are incompetent makes sense. You can apply this logic applies to all functions.
On the other hand, at merely single-digit servers, I'm guessing the fixed cost of personnel referenced in the GP comment would dominate.
But most firms don’t know how to employ a you to show them this way.
This article only looks at "seen" costs, and assumes that there are no "unseen" costs to running on-prem. Many companies do not have the operational maturity to run on-prem well. The result: high cost of operations, low availability, and large increase in time-to-value.
Second unseen cost: everybody becomes their own SI. So far nobody really sells the "whole stack" for running on prem. I mean hardware, network, virtualization, application, traffic mgmt, etc. I have to buy stuff from two dozen different vendors and cobble it together into a high-labor, rickety Jenga tower of stuff.
If a company was going to run their own data centers, they would presumably hire someone that knows what they are doing to lead the effort instead of trying to do it by trial and error.
Yes, you have to plan, yes you have long lead times to get new gear, DC space, etc. But once it's running, failure rates are pretty low. At least in the single digit thousands and servers.
I've had to deal with much more frequent and odd types of failures on cloud infra than with on-prem.
Supply chain management and coordination is, IMHO, the most difficult part of it and is often overlooked in these discussions. Herding vendors to deliver hardware on your timeline from overseas supply chains doesn't always turn out the way you want, even with the best planning.
Isn't that the target audience for this sort of change?
I mean, do you see teams of networking , siteops, high availability, and security engineers hanging around doing nothing and just waiting for a company to decide to go in-house?
No, because those teams do not exist in free-range. That's something a company needs to build and train and experiment from the start until they are able to learn all the lessons.
My company (which is in fact part of the charts in the article) had every single metric across the board spike 15-20x basically overnight when the pandemic started last year in March. Our entire infrastructure burden was clicking a few buttons on the AWS console and making sure everything was provisioning and scaling as needed. If we had to send out people to buy hard drives and server racks at that time, there is no chance we would have been able to meet the extra demand.
Plus, if you give me a few dozen capable engineers today I'm not going to waste their efforts on rebuilding AWS to get a best-case few percentage point return on our cloud spend. I'll launch a new product for our customers instead.
I’ve also worked for a place so large in AWS we hit walls where we hit Amazon’s literal physical limit and had wait for them to go out and buy the hardware for us to provision.
For Dropbox specifically I don’t see them being any form of special case, if it saves money and they obviously had operational experience so it makes sense
The implication here being, "Well, if Dropbox can save this much, then think about how much everyone else cans save!" But in fact the opposite is true. Dropbox sells disk storage on the cloud. For them to do so by effectively reselling someone else's disk-storage-on-the-cloud platform is obviously not high margin, and they'd be better off building it themselves. So, sure. Anyone else also offering, disk storage on the cloud, or compute clusters on the cloud, or otherwise just reselling someone else's product on the cloud, will probably have higher margins doing it themselves. But that is certainly not most companies. So, yeah, it's "just Dropbox".
Disclaimer: I work for a cloud provider, but these opinions are my own.
While repatriation can make sense at a larger scale company, startups and SMBs can yield the same benefits discussed in this post by simply tracking and optimizing cloud spend.
We try to make this as easy as possible for people with https://www.vantage.sh/ - where we're already helping thousands of individuals, startups, SMBs and enterprises as it relates to AWS.
The work was very interesting, because a lot of it was actually building a private cloud for internal customers & the work primarily centered around a virtualized data-center aimed at boom-bust cycles of games (15+ million users for 6 weeks, drop to 2 million for a month and down to a million in another week).
The issue is that the infrastructure cost is somewhat constant when dealing with that sort of fluctuations in revenues, so the cost to revenue ratio was unpredictable (while the cost was).
So what happened in the end looked a lot like a fire-sale of hardware when the cost was unbearable, while if it was an end-user cloud, that low-demand phase would be able to cut losses as a spot instance or something.
Anyway, a few years after I left, back to EC2 it is[1].
It took me about a day to scope out a BOM that would well surpass that (maybe a year's expected growth) of whitebox server gear, and get some quotes from nearby co-los.
IIRC hardware capex was about £15,000, and monthly rack + network something around £3,000. We relocated within a few weeks.
One small bonus was predictable & consistent performance -- back then, the EU-west EC2 offerings were extremely sensitive to noisy neighbours, and our benchmarks never gave anywhere near the same results twice.
I think for genuinely elastic loads, or if you're really addicted to some vendor-only services, it certainly makes sense. I suspect most customers are overly optimistic about how elastic their requirements are, and their stack can be.
The salaries and hardware cost was paid for even if the games had a bust.
The period where it worked well, the games had boom-bust in somewhat controlled fashion where farmville -> cityville -> frontierville -> fishville etc, the traffic would move around rather than die down entirely.
The world turned mobile-heavy and that whole pipeline fell apart while they were restructuring into mobile games (words with friends etc), when the hardware had to be sold or the payments would start to hurt.
If you self-manage, your capital investment is initially higher, and lower over time. At the same time, the effort it takes to reach the same results is always higher, and the quality of the end product may be lower, depending on how much service quality affects your product.
If you pay for managed services, your initial investment is lower, and higher over time. But at the same time, you require less effort, and you get higher quality outcomes.
This is obvious to anyone who has worked in the industry and done both. First, host your own service: JFrog Artifactory, Atlassian Confluence, GitLab, whatever. Now rapidly increase the demands on this service. As demand rises, quality will decrease, because it takes a lot of time, effort and expertise to build a very reliable hosted service. Now switch to a managed cloud instance. Suddenly, the service's average quality increases. Performance is steady regardless of increase in use. There are virtually no interruptions to your product or development.
The impact of a service's quality and reliability has ripple effects. If poor service quality slows down development, that means development quality will go down as people cut corners to try and meet deadlines. If the service is used for production, it means your product's quality will suffer, and that effects your bottom line. So a huge amount of the actual cost is not just paying for a service, but also how much business value is generated or lost due to service quality.
There is simply no way to replicate a managed service without becoming a managed service provider yourself. You have to become a whole new business within a business. It's like a yogurt company also becoming a dairy farm. Running a farm is not easy, and you will screw it up for several years. Seems obvious for farming, but for some reason people always underestimate this when it comes to technology.
On paper, the Cloud's value proposition is scalability. But in practice, the true value is actually as a force-multiplier for your product's quality, reliability, and time to market. (Time to market not just being "I launched my startup" but also "I released this new feature before my competitor")
Avoid doing undifferentiated work. Outsource it when you can. Focus on the core product. I think this is especially true as you're scaling a company in the current SWE labor environment. If your business is growing quickly, you probably have money, and can get more funding when you need it. But hiring more SWEs is always a challenge. That is your scarce resource. Don't waste it on undifferentiated work.
All that said, at larger scale and more mature businesses. I think the core thesis of the article makes sense. You may be in the slower growth part of the S curve, but you can still increase your company value and profitability substantially by increasing margins. One way do do this is to invest in cheaper infrastructure.
Then you're taking on some undifferentiated work, but in service of better margins.
I've worked for a handful of very large businesses (10s of thousands of employees, profits in the billions) running infrastructure and services. With one exception, they all sucked at managing their own hosting. The only one that did well literally poached all the core employees of a web hosting company, and then listened to them.
Except for that one company, none of the others was willing to build their own internal managed hosting company. They all didn't staff right, didn't fund right, didn't do support right, and their org structure and financing model had no capacity for a single department to straddle every BU. (Well, ok, "IT" often did, but "IT" wasn't directly managing and running the underpinnings of every product in the company)
In fact, some of the largest, most well known companies in the US actually moved to the cloud specifically because they sucked so bad at running their own shop that they needed to stem the bleeding by just buying Cloud services that worked. And that's after they hired consulting/management companies to try to jointly manage their self-built setups.
The article makes sense. The idea that larger businesses can figure out how to do it right at scale, sounds logical. But my whole career is one long proof that those theories don't match reality. In fact, I would go so far to say that the larger the organization is, the harder it is to self-manage. The only general category of company that I think is OK to self-manage hosting is one where either almost all of the profit is coming from keeping hosting costs down, or the business value is not impacted by hosting quality at all. I think that's a very small proportion of tech businesses today.
If you have moved to AWS / other then that time is around 5 minutes.
If you are in a major fortune 500 and need a new server, quite often that time will measure in months (yes really).
This simple equation just blows every other cost/benefit calculation out of the water.
I may have missed it in the discussion.
Last time I was procuring racks of servers, it was about a 6 week lead time, sometimes more. So yeah, you have to do a lot more capacity planning, and probably keep capacity in reserve for things that are unexpected.
I've had to go around to every engineering team and try to get them to tell me how much gear they might need in the next 1-2 quarters. They rarely know, so it's a lot of guess work.
So in a situation with a quickly growing product, that grows in an unpredictable way. Or a really new product where you just have no idea what the adoption is going to look like.
Then yeah, getting those new instances right now is really really important. At Segment we very often needed more instances in a huge hurry due to some customer load spike.
And sometimes AWS ran out of the instances we needed in the region we were in. So we had to get creative to use other instance types.
Also the cloud provider pricing structure heavily incentives you to actually do some capacity planning. On AWS, Savings Plan gives you huge savings.
But if your workload is fairly predictable, then this on demand benefit isn't nearly as compelling and at the right scale, it makes sense to start buying servers to put in your own datacenter.
Not everyone needs a server immediately. And not every use case will generate business value by having that server immediately available. This blog post cuts to the core of that: cloud gets you from product market fit to scale, but once at scale, it’s time to revisit capex vs opex.
Outsourcing the "where the hell is my container full of hard drives" phone calls, and subsequent scramble when you find out it was delayed for four weeks but they failed to notify you, is almost worth it. (In fairness, this mostly goes away if you select and manage your vendor relationships well.)
Now if someone wants 20,000 additional instances overnight that’s going to be a problem, but you can certainly get enough to prototype and and run a small scale rollout while your capacity paperwork goes through.
Yes, but the benefit of cloud is we only work to optimize those which gain market adoption. For every twilio there may be 100-1000 startups that did not make it, it's a good thing if those were constructed rapidly, tested for product market fit and then turned off without optimizations applied.
Also if cloud adoption actually accelerates building, then it may also aid its adopters if there is a race to product market fit.
The problem with going on-prem is that for most companies this is so complex and hard to operationalize that they will likely just fail at that attempt.
Cloud biggest contribution besides seamless infrastructure access, is the way it has lower the overall technical competency threshold to operate an internet business. It's really easier to find someone that knows AWS or Azure than is to find that someone who knows how to build on-prem infrastructure.
I imagine that repatriation for the average large cloud business will only be possible for those who are already operating in modern infrastructure constructs like Kubernetes.
If you have critical workflows build on propietary cloud services, you're probably fucked. You will have to rebuild every single service from scratch and make it scalable out of the box. If you're a multi-dimensional business that sounds like a nightmare.
I honestly think that Dropbox was capable of doing it just because the nature of their business.
At the primitive level you're mostly replacing your storage vendor. It happens to be that storage is their business model so it's quite evident that repatriating that critical piece of your business will yield great gains. I don't think it's that simple for other cloud businesses.
Whenever a rewrite happens in software there are usually massive performance gains. The same thing happens with infrastructure.
> So what can companies do to free themselves from this paradox? As mentioned, we’re not making a case for repatriation one way or the other; rather, we’re pointing out that infrastructure spend should be a first-class metric. What do we mean by this? That companies need to optimize early, often, and, sometimes, also outside the cloud. When you’re building a company at scale, there’s little room for religious dogma.
- HW that fails, go get a new one, cost + time
- backup and geographical redundancy
- security
- certifications
- tooling to manage and forecast
- black friday! scale for few days ... more HW?
- etc ...
It is way more convenient for the average enterprise to pay premium and be done with it. So they can focus on the actual business and not building infrastructure and tooling to manage it.
If you go cloud, you get a handful of very large suppliers, that provide a lot of non-standardized services that are probably going to lock you in like hell (as the article wel says).
If you build infra, you're using /mostly/ commoditized supplies and skills.
If the cloud offerings where interchangeable and the industry reasonably fragmented, the margins of cloud providers would be slimmer and the paradox would probably go away (in favor of: always cloud!).
See for example Porter's analysis framework [1] and how your positioning changes in the two cases.
https://en.wikipedia.org/wiki/Porter%27s_five_forces_analysi...
Given that interactions to the cloud are through a relatively small set of known APIs, I'm surprised that there isn't there a service which replicates the APIs but with the endpoints being on "repatriated" hardware.
Designing and implementing a highly available and resilient storage solution on-prem is much more harder task that will take 99% of time & budget to do properly compared to a compatibility layer.
I’m looking for a mechanism which apes existing cloud APIs directly.
But there's also: https://www.eucalyptus.cloud/
Little Miss Muffet
Sits on her tuffet
Watching her billing grow
Her devs ever hectoring
On the need for refactoring
But what do these uni-brows know?
Another thing that came to mind reading this was - what would I need to make such a switch (from cloud to your own data centre)? And then I thought, you could provide a business doing this - commoditised hardware and services, providing the real-estate and the engineers, and more importantly, providing the same kind of service stack you find in the cloud. Then I thought, that could be a cool business for someone. And then I thought - there's nothing stopping Amazon or Microsoft or Google from offering that.
I guess if you have a very clear understanding of your needs, a mature appreciation for what it will take, and a good handle on your likely growth, then taking control of it all may make sense. But for a lot of business (especially those not actually tech businesses per se), it would more likely be a return to pain.
The same is happening with SaaS. SaaS eliminates the need for in-house IT! Except it doesn't. It just means you now have a bunch of recurring SaaS costs that you are locked into forever because they have your data and you still need IT people to babysit your massive cloud/SaaS sprawl.
Not following fads and buzz is a huge competitive advantage in this industry. Founders and chief engineers / CTOs / CIOs take note. Just make sure you can explain why you are not using (insert latest buzzword here).
The bottom line is that you should analyze the situation using your work load, your numbers, your culture, etc., and decide what works the best. Sometimes that's managed cloud. Sometimes it's unmanaged cloud. Sometimes it's bare metal. Sometimes it's on-prem. Your mileage will vary.
> Cloud became the fad, and the buzz was that cloud saves money, so everyone goes cloud
Cloud services _do_ save money for businesses of the appropriate size (meaning, those that can't or shouldn't be focusing on physical servers and networked hardware).
> The same is happening with SaaS. SaaS eliminates the need for in-house IT!
Again, SaaS _does_ eliminate the need for in-house IT _for certain classes of businesses_.
> It just means you now have a bunch of recurring SaaS costs that you are locked into forever because they have your data and you still need IT people to babysit your massive cloud/SaaS sprawl.
The cost of employing someone to babysit a SaaS product is much lower than the cost to employ someone to build and maintain an equivalent in-house SaaS product. If you're fine with vendor lock you're saving money.
> Not following fads and buzz is a huge competitive advantage in this industry.
Ignoring all nuance and claiming things that are popular are "just fads" is, to me, so much of a competitive disadvantage that it almost certainly outweights any perceived gain.
You pay a higher price for what you use than if you maintain those tools yourself... but you don't have to maintain those tools yourself. It's hardly a buzzword, it's almost defacto standard nowadays.
1) It takes amazing intestinal fortitude to fight back against the fad tide. And, when things go wrong, your decisions get the blame even if they aren't at fault.
2) As the article points out, cloud is FINE for startups and small companies. If my startup company reaches gigabucks in revenue and I now have to worry about $75 million in cloud spend, I've done my job. And then some.
3) Cloud is often about blame and liability transfer. I don't want the company website getting hacked to be my problem--I want it to be somebody else's problem. I'm willing to pay for that.
Public clouds, like AWS, have cut their storage costs by more than 50% since Dropbox built their own infra in 2015.
The last time I ran physical infra, we had sever racks that I'd originally bought, running in production 5+ years later.
Once you're past the depreciation schedule, they are basically free except for power and space. Yeah, at some point and scale, it can make sense to get new gear that is more power and space efficient to pack more into the same power and cooling budget.
AWS still lets me launch c1-c6 instance types. So they still have the older generations sticking around. Yeah, the newer ones are usually more cost effective, but you do have to do work to migrate to them.
Exciting times for decentralized storage!
When you're small, outsource.
Should the company find traction, and as you gain size, the arguments in favor of outsourcing fade, and bringing the IT in-house makes sense.
> tie the pain directly to the folks who can fix the problem
That's why we're building https://github.com/infracost/infracost for engineering teams (free open source)
Kubernetes only makes sense from Google's perspective to sell you the managed version "from the creators of Kubernetes!" once you get tired of trying to wrestle it into submission.
K8s isn't that complex from a cluster operators perspective if you've managed bare metal Linux systems before. Set up some PXE + auto provisioning and you can make your node enrollment process:
1. Buy hardware from an OEM.
2. OEM ships you a piece of hardware.
3. Remote hands plug in hardware.
4. Machine PXEs, boots, and is turned into a worker on your cluster.
Canonical (MaaS), VMware (vSphere), and Rancher Labs (Rancher) all have products aimed at running Kube on metal in your own datacenter.
So, essentially, if you start your startup off using Kube in a managed cluster you can, at some point in time later, directly move your code off the cloud and onto your own machines for cost saving. You just need to figure out how to run your DBs and LBs but that's a solved problem as mentioned above.
k8s lets us move fast - we can easily deploy new services in no time. We can also run the entire stack on our desktops in k8s, or have our laptops be a node in a mostly cloud k8s dev deployment and debug services locally.
That completely overlooks the actual benefit and the reasons driving cloud adoption, which he states early on and then fails to integrate here. The real cost/benefit analysis would take into account the amount of time, money, and opportunity cost saved by using a cloud in development, a critical time that determines whether there'll actually be a profitable company eventually. Optimizing the cloud costs is certainly important, but being able to spin up highly integrated systems on demand offsets a huge amount of time and capital during development.
A takeaway might be that, e.g., AWS, should offer even larger discounts for large scale operations in order to retain mature customers.