Source: I work at Lyft.
Worked for us at Twitch.
Their surges are nothing like yours.
If you have engineers making less than $250k annually, then you have a lot more staff. $8mm monthly is a LOT...
Amazon's clearly making a profit after $8mm monthly.
Lyft could definitely build and maintain their own infrastructure for this kind of money... probably do it better (customized to their needs) and cheaper.
Businesses don't flagrantly throw around money just to upset people. There are huge advantages to offloading non-primary business costs to other businesses.
Netflix is doing this too. I think we can assume not all of them are just idiots that haven't figured out they could build this themselves.
But businesses do throw around money for the wrong reasons, and keep on doing so if that's the status quo. No one gets fired for buying IBM.
> Netflix is doing this too.
IBM stuff was bought by a lot of people.
> I think we can assume not all of them are just idiots that haven't figured out they could build this themselves.
That statement is very misguided and misses the problem. For example if you built your infrastructure around a specific solution then you also end up building a team of professionals whose livelihood is tied to a specific supplier of said infrastructure.
Businesses are wasteful because that's the natural status of a bureaucracy. They aren't throwing away money on infrastructure because they are unaware, they are spending more than they potentially have to because infrastructure isn't their core business.
> IBM stuff was bought by a lot of people.
That's such a tired argument. Just because they could save money doesn't mean it's a good idea, and with Enterprise pricing from Amazon combined with tax advantages, you honestly have no idea how much "cheaper" it really is.
> That statement is very misguided and misses the problem. For example if you built your infrastructure around a specific solution then you also end up building a team of professionals whose livelihood is tied to a specific supplier of said infrastructure.
No, the fact that you think this is a "problem" is the problem. Do you honestly think dev ops guys couldn't figure out how to use a different tool? By your own logic, you also shouldn't build data centers because you end up building a team of professionals whose livelihood is tied to managing your own infrastructure.
That's not true at all. The "isn't their core business" argument is meaningless and absurd. Any company, big or small, does not want to waste 300M dollars on something they don't need, whether it's their core business or not, particularly when said company is still far from turning a profit.
> That's such a tired argument. Just because they could save money doesn't mean it's a good idea
You are aware you're stating that baseless assertion on a discussion on how a company which is burning through cash and looking for investors is needlessly wasting 300M on infrastructure costs.
> No, the fact that you think this is a "problem" is the problem.
Needlessly spending 300M dollars is a problem in every single business in any corner of the world. I have a hard time understanding how someone can throw around the baseless assertion that this sort of inefficient while operating at this particular scale is not a problem, and pointing out this problem... is the problem? That's crazy.
It feels like you could fit half a Lyft into live low-latency transcoding and redistribution of just the top 10 streamers feeds on Twitch.
Source: also work at Twitch.
Oh, or are you making fun of using Bare metal?
I think the parent was pretty clearly sarcastic and suggesting that this was a bad idea.
The problem is that provisioning, reliability, and security are by themselves really tough problems. If those issues aren't in your company's core competencies, it's not necessarily efficient to invest in building out all of that.
I look at it as the question: can you get the same set of agility/reliability/security guarantees for your narrower set of use cases by paying for your own hardware and engineering? I won't even begin to pretend I have any answers there, but I think that's the calculus.
Maybe that's just the story cloud providers tell you.
Until you try, do you really know if it's all that complicated? People have been running datacenters for a long time, and not all of them work for Amazon.
But there may be also a beneficial side effect of having gearheads around, and maybe that's the real cost to going cloud.
Since we're internal and we manage a lot of capacity, we do often provision and roll our own equivalents of things that cloud providers will sell you, rather than just buying a cloud solution. It's often ambiguous whether it was a good use of time/money. If it weren't for the economies of scale that kick in at the sheer size of this operation, it would definitely not be worth it.
- Recently had to purchase new servers, because of signed contracts the only servers we were allowed to purchase and put in the datacenter were four years old and technically EOF.
- Firewall changes, AD changes, provisioning a VM, etc. are 48 hour turnaround. Purchasing new hardware requires 4-6 weeks.
- Had an intermittent issue with their edge firewall, it'd slow certain connections to a crawl and eventually they'd timeout. Took six months to fix it, for the first three months they told us it wasn't their fault (turning off their deep packet inspection ended up fixing it). I still remember when we opened the first ticket about it, and the reply was "no other customers are experiencing problems" and it was closed.
That's just a few examples of how painful it can be. To give you the other side of the coin, having worked with an enterprise contract in AWS, we were having an intermittent issue with DNS resolving failing for a few seconds every few days. They put an engineer on it full time till they found the problem (we misconfigured it), and it didn't cost us anything more than the enterprise support. I was actually shocked they'd invest that much on such a vague issue.
Yes AWS is expensive, but you're getting world class engineering proven at scale, and access to some very smart/motivated people to support it (and they have access to the teams who built it, when they can't solve it). I don't think I'd ever choose managed datacenter over AWS/GCP/Azure/etc. Either do it in-house where there's accountability, or use cloud providers who have proven they're competency.
To be clear, I'm talking about VPC/EC2/etc. I can't really discuss a lot of their higher level and newer managed services; they either weren't as good, or I haven't tried them. But the bedrock these clouds are built on is solid, and that's worth paying good money for.
> I don't think I'd ever choose managed datacenter over AWS/GCP/Azure/etc.
Who mentioned managed datacenters? I'm pretty sure people are talking about leasing space and doing everything else in-house.
- storage clusters
- database clusters
- compute clusters
They are often very easy to setup, but when things go wrong, they go very wrong. And welcome to a stressful environment because if you can't figure it out and your people can't, well, your business just sits and burns while you do.
Even when AWS has a system-wide outage, it's nice to know that I don't have to be dealing with those underlying problems anymore and I know they have the best people working on them.
I cannot put into words, after operating MySQL clusters on my own and playing back transactions after failures, how nice it is to use AWS RDS and how it's just been zero problems. Zero. I sleep through automatic updates of our database system with RDS. I would have never done that on our own system.
And in most places, even "managed" leased hardware, you still will need to purchase/lease and run your own hardware firewalls and ddos mitigation. The datacenter might offer that protection "built-in" but you'll soon find the limitations of that offering when you face a substantial attack.
Having spent my entire working life automating infrastructure of all kinds I know you can achieve an enourmous increase in efficency rather easiliy with a few well placed automated processes.
I’ve always been baffled by the fact that at any given larger company there are 100’s of employees trying to supply the business with tools to automate business processes — the IT dept.
Yet, they are completely incapable of using these very same tools to automate their own ”business”. And the resistance I’ve been met with at different places through the years when trying to implement the simplest of automation is massive.
I used to laugh at the ”cloud” bacause, back then, at 25 years of age, sitting at a medium size company with boatloads of cash, I assumed everyone was doing it the way we were; automating all the things.
Now, many years later I’v obviously realized that many places simply does not have the right culture and mindset as it’s not “core business”.
I believe however that this is changing, and changing quickly. In many ways thanks to the “cloud”.
Operating bare metal at scale requires talent that doesn't exist, not necessarily at an engineering level, but at all levels.
As an example, I worked at a place that had a large bare metal deployment, i.e. >1MW worth of compute. It was woefully inefficient and costly to operate. The product that they offered required network QOS and compute with real time capabilities, neither of which was available from any cloud provider at the time.
One of our executives (formerly a leader in the DC ops org at AWS) left the company to be replaced by another executive by another well-known silicon valley org who then insisted we should migrate everything to the cloud.
I showed him the relatively easy math that efficiently utilized bare metal was way less costly and that the aforementioned QOS and RT requirements would be a deal breaker anyway. He failed to fully grok this and remained insistent. When I quite, he seemed surprised. After the fact, I discovered that they'd made a deal with IBM to move everything into their cloud. A year later it was an utter failure and they abandoned the project.
There are lots of folks in the valley with lots of experience on their resumes that suggests that they should be capable of understanding these kinds of things that simply don't. Lacking that understanding leads to poor decision-making, which leads to failure, which leads to risk-aversion, which leads to everyone believing that it must be cheaper in the cloud.
Or so goes the old adage, "nobody ever got fired for buying IBM."
EDIT: To whoever downvoted this, the commenter hasn't listed an email address, or I would have reached out directly. This is an honest attempt at communication that doesn't require someone to break anonymity.
Though the question I received was somewhat nonsensical, which was to be expected.
RDS doesn't really scale without costing a fortune. It buys you HA and backups. Great, but what if you need performance?
DynamoDB? It scales in terms of IOPS, but again, it's unaffordable.
SNS exists and isn't terrible, but why wouldn't I just run Kafka?
But if you need bleeding-edge Postgres performance, you hire a DBA, and they probably build something on EC2 or bare metal.
———
As I understand it, RabbitMQ is probably a better point of comparison for SNS/SQS, and Kinesis is the Kafka peer.
Regardless, the reason you don’t “just” run Kafka is: you don’t have a team that knows how to tune, deploy, and operate a production Kafka cluster. I learned enough about SNS and SQS to get it running in an afternoon, and I really haven’t needed to think about it since. Kafka (or RabbitMQ, or ActiveMQ, or etc) need instrumentation and monitoring and patching and quorums and capacity planning and etc, and at some scale those are worthwhile, but that scale is MUCH larger than what most Kafka clusters are actually serving.
———
The theme here is: if you have a business requirement for 90th percentile specialized performance, great! Hire domain specialists who can make your systems run at that tier! But for everyone else in the world, when you can get usage-based pricing, elastic resources, and automatic durability and patching... why would you go to the trouble of learning how to deploy and manage a service?
Perfectly willing to admit I'm wrong if and when that time comes. At this point, that's my theory.
Bare metal works when your workload is well-defined and understood. Then you can actually put reasonable estimates for what you need and hire/purchase infra accordingly.
The balance here is tricky. Based on public data, it seems that Netflix has ~$16B in revenue against $300m/yr cloud spend. 2% seems much more reasonable to me.
I feel like a drive toward efficiency is a worthwhile endeavor for a startup in terms of establishing a competitive advantage.
I remember how hard it was to hire senior operations people. There are not many of them, and there are not many of them at the level of being able to deliver something amazing. The ubiquity of the cloud has only made these kind of experts less common.
Every place I've worked that did bare metal was always drowning in maintenance instead of working on the next big thing. And no big surprise, our internal infrastructure was nowhere near as high quality or capable as AWS. And most of our developers had experience working directly with cloud providers, without ops people, so we were delivering them a worse experience and slowing them down, and we required more ops people to help them and maintain it and keep everything online.
Also, a move to IBM's cloud isn't the greatest example. I had hundreds of bare metal servers in an IBM-owned datacenter and their cloud offering was consistently behind AWS/GCP; if anyone recommended IBM cloud to me I would have laughed at them. It seemed to me that IBM was trying to up-sell on the "cloud" buzz word without actually delivering anything except higher prices, just like how they're now trying to ride the buzz of the blockchain.
Dropbox is a good example of a company that took quite a while to move to their own platform, away from AWS (and they still have 10% of their stuff in AWS to this day). Dropbox is basically a storage infrastructure company, unlike Lyft, but it still took them years to invest in the development (and migration) of that custom platform to replace AWS, an investment that not many companies are going to want to gamble on, especially if their primary business is not storage:
https://techcrunch.com/2017/09/15/why-dropbox-decided-to-dro...
And I think it's telling that Dropbox started on AWS, grew the business on AWS, and moved to a custom platform once their business model was perfected and they wanted to cut costs prior to going public. If Dropbox had started on bare metal from day one, would they have been able to pull it off?
There's nothing you've written that I disagree with. It's easy to do the math that shows where bare metal saves money inclusive of the labor costs. For some reason most everyone seems to fail at it. I could expound one why, but this:
>I remember how hard it was to hire senior operations people. There are not many of them, and there are not many of them at the level of being able to deliver something amazing. The ubiquity of the cloud has only made these kind of experts less common.
Those folks just don't exist. Building infra is more than just buying infra. It takes actual development, which is why I think so many fail at it.
Your anecdote about Dropbox is telling. They adopted cloud, and more importantly cloud methodologies and then went back to bare metal. There are others that have done the same. I recall a talk at an Openstack conference given by Verizon in which they described their approach. Developers begin in AWS, utilize a cloud-based approach, and then when cost concerns become an issue, they aim to offer similar services in-house on bare-metal.
This is true but its never really hit me before even though I've already been operating based on the assumption that trusting the cloud is less risky than trusting my own skills.
It slowed down both them.
Remember that staff cost money too!
Conversely, with GCP 4 years ago now had some support issues - didn't come away impressed - I'm convinced even internally GCP isn't well doc'd or something.
But what I paid for and got on aws support is so far out of whack there is NO way they made money on my account for that whole year. And the person was actually competant which was a shock. So many "technical support" folks seem like idiots.
Comcast for example, I'd purchased my modem, they started charging a rental fee - I had to call these bozos every month to reverse the charge - a total waste of time. I cancelled finally - I just couldn't take it, and each one lied to me or didn't have a clue. Things like condescendingly saying - you have to pay for the modem.
They want to make money brokering rides.
Taking on their own cloud infrastructure -- in theory -- could economically make sense. But that's just an extra layer of risk and complexity they'd rather forego to focus on their core business.
After all, their core business is already losing $930M on $2B in revenue. They're cash-flow doesn't put them in a good position to make large up-front investments on data centers.
So, yeah, like a broke renter in an expensive city. In theory, it might be better to buy a house, but you don't have the down payment, and maybe you should be focused on increasing your earning power rather than saving money anyway...
That's not a judgement on whether it's worth it for Lyft or not, but especially for a growing company with spiky load the decision is not just a dollars to dollars comparison.
If you gave me $300m to spend, largely up-front, for significant capex purchases? Sure. We could do it. The team I would build would also probably still make mistakes that AWS et al have already largely learned how to avoid, but we could do it. But capex and opex are very different beasts. By the end of that three years I'm already looking at spending way more to refresh what I bought at the start of that three year period because I'm starting to near the end of early contracts and I'm figuring out how best to wrangle, in a way that makes the rest of the business succeed most optimally, a now-heterogeneous environment, etcetera etcetera and etcetera. It's all solvable. But whether it's cheaper, at scale, and more reliable, and presents a unified tool for use by the business...that's a harder question.
Understanding how capex and opex work and how they differ is pretty critical to successfully running an engineering organization, to say nothing of a company.
The reason that AWS, Google, Azure, et.al do so well is that they don't just buy some servers. They do actual capacity maangement, and not a very good job of it I might add. They also manage the lifecycle of every component in the infrastructure such that the next iteration of that component is understood and interchangeable.
Network architecture, for example, should suit the needs of the application, but should also be decoupled from the underlying hardware as that hardware is going to evolve.
Compute is fairly straightforward as well. At the data center level, one makes a bunch of 400W holes. What you fill those 400W holes with is relatively irrelevant.
The care and feeding of fleets of (physical) machines is really, really hard and not to be underestimated.
If you do that, though, my answer will be "right, so we're done here."
Considering how most layman are completely wrong in their understanding of finance, I'd say that isn't a good endorsement...
It's all about leadership. The dearth of skilled leadership is the issue. I'd wager this is how some FAANG companies are managing this. They're hiring people that know what they're doing. One doesn't need to design and build their own servers and network hardware to do well at the scale of folks like Dropbox or Lyft.
Cloud adoption is all about making the issue someone else's problem, which is only kicking the can down the road. Eventually, every company that does a thing will realize that their survival is contingent upon becoming a software company that does that thing.
If you have $100m OpEx per annum, it'll cost you maybe a point or two to convert that to $300m CapEx.
- person who knows how hard it is to run your own infrastructure
- person who knows how hard it is to run your own infrastructure.
> fierro Profile: SWE @ Google Resource & Capacity Planning
I think you mean "Person who's job it is to convince others it's really hard and they should just buy your product"...?
There is way, way, way, way, way more to a running a successful business than "saving money".
Most are running a few small internal-facing servers hosting some internally developed apps, and need very little resources.
Just run ESXi, XenServer, Xen or something, and spin up a few VM's on a few thousand dollars of hardware, get a couple people to maintain it, and be done.
Even at large scales, like Lyft, having your own internal team and hardware is going to save money. Amazon is profiting off your instances... which leaves room for you to do it for less. Maybe not $7mm less monthly, but even a $1mm savings is significant... but likely a lot more.
Fast forward to today, and now it would be a serious undertaking with serious risks to move off AWS, not to mention the costs of building up the staff and assets to reimplement their requirements in parallel of AWS until reasonably confident they can flip the switch and still have an operating company afterward.
So, they're probably stuck - beholden to Amazon's whims and pricing mood of the day. They've bought convenience from Amazon in trade for massive technical debt, one which may be even more costly to get out of... Or impossible.
AWS isn't going to get any cheaper in the future..
Those free AWS credits Amazon gives students really pay dividends.
Arguments like yours are why business people tend to roll their eyes and ignore engineers when it comes to anything outside of engineering.
Not trying to be dismissive, but you are so far from the mark I don’t know where to start...
I agree that it would be a serious undertaking to move off AWS today. But it's probably also going to provide marginal benefit. No one on the finance side of their business is probably losing sleep over it. If/once it makes sense to move off then the finance dept will tell the eng dept they need to reign in infrastructure cost...and eng will do that.
The “whims” of Amazon’s pricing are no more unpredictable than the pricing “mood” of your colo or your hardware vendor.
So basically, I disagree. These aren't estimates.
I was a cloud skeptic and ran Tech Ops (including our DCs) for years. About 5 years ago, it dawned on me that even owning the whole budget for Tech Ops, that I wasn’t capturing the full costs of trapping my org onto our in-house solutions.
At tiny, small, and medium scale, cloud is obviously the way to go, IMO. At large and huge scale, I think letting some hybrid leak in where systems change rarely and cloud costs are WAY out of line (DropBox storage, Netflix CDN, etc) makes sense.
And management of such organizations might also have ideological tendencies that further skew the calculation.
There are certainly examples for big companies that benefit from having their own infrastructure (i.e. Dropbox since they have relatively specialized hardware needs compared to what cloud providers set prices around), but the number of people you need to hire to build and maintain datacenters is very high.
See https://blog.twitter.com/engineering/en_us/topics/infrastruc....
* If the hammer manufacturer decides not to sell you any, you'll still have hammers.
* If the hammer manufacturer gains enough power to fix prices, you won't be paying them exorbitant prices.
* If the hammer manufacturer or their country gets embargoed and you're unable to legally purchase their hammers, you'll still have hammers.
All the above grant you a strategic advantage since you'll still have the necessary tools to continue your business while your competitors won't (or will have to pay much higher prices for their supply of hammers).
Sure. If they are paying $100M/y on hammers, it's at least worth running the numbers and investigate alternatives.
Strategically, you probably want to focus on what your core competencies are, even if you could in theory do something for cheaper. It's easy to ignore the foregone best alternative of iterating on your own product instead.
This is expensive and risky and also difficult to do in piece meal
Disclaimer: former AWS + Amazon employee