Uber migrates microservices to multi-cloud platform running Kubernetes and Mesos
uber.com
uber.com
This is something that most companies don’t do when they say they want to do $x to “prevent lock in”.
Uber actually is testing for portability along the way.
One of which is putting a price on HEAD requests in S3 (?).
AWS already gives long term price discounts/guaranteed prices for reserve pricing and Big customers already have negotiated contracts.
Even if you have done everything in a “cloud agnostic” way, “infrastructure has weight”. Any large migration isn’t just technical , it involves project management, organization training, regression testing, compliance testing, security testing, architecture review boards, vendor negotiations, firewall changes, coordination with third parties who may only allow list certain IP addresses, data migration, etc.
Heck they often have multiple physical network connections to the cloud provider (Direct Connect)
Anyone who thinks they can run everything on K8s and they have “cloud agnosticism” has never done a very large scale migration.
You would be amazed how long it takes to do a bog standard lift and shift of a hundreds of plain jane VMs and VM hosted databases. You can’t get anymore cloud agnostic than that.
source: I’ve done a few over the years in both the “real world” and working in the cloud consulting department at AWS (Professional Services). I no longer work at AWS and have no specific loyalty to AWS.
Egress price makes it worth migrating away from those three.
If you choose AWS (or Azure) and a region goes down - everyone else is down too. “No one ever got fired for choosing IBM”.
Choosing the most popular vendor - AWS, Salesforce, ServiceNow, or whatever vendor is in the upper right Gartner magic square quadrant never gets questioned by the powers that be.
And for you to just say “it’s okay to be down an entire day” because of egress cost tells me that you have never done infrastructure requirements analysis at scale.
First you have to assess the cost of being down for a period of time, then you have to access RTO, RPO requirements and not all workloads have high egress costs - especially things like data lakes that may have a lot of ingress and processing costs, but relatively low egress costs.
I’ve done a lot of different cloud projects over the years from lift and shifts, to data lakes, to cloud call centers, to serverless, to ETL jobs, you can’t just blindly repeat “egress costs” in a vacuum without understanding use cases.
> Egress price makes it worth migrating away from those three.
> Even if the alternate cloud provider goes offline for an entire day it still would be worth it financially compared to AWS because egress is so expensive there.
You never qualified either with “in my particular use case”. If you had, I would have had no argument. I haven’t been flown into your company along with SAs, sales, project managers, etc for a week to do a proper “as-is” assessment and to see what your requirements are.
I haven’t accessed the competencies of your staff or determined what is your competitive advantage and what is the “undifferentiated heavy lifting” in your company.
I would never make any blanket statements without knowing your specific use case and automatically assume “cloud” is always the right or wrong answer
This suggests you are looking for a single example where the pricing of the big 3 is compratively high compared to the competition at a point where it worth it to switch. I gave the example that the price of egress is one cost which is not competitive. If I had instead said that SQS was not competitive obviously that wouldn't matter to businesses that don't use it enough to make a difference.
Microsoft and AWS have versions of the “Cloud Adoption Framework”
https://learn.microsoft.com/en-us/azure/cloud-adoption-frame...
https://aws.amazon.com/cloud-adoption-framework/
And the TOGAF framework has something similar
https://pubs.opengroup.org/architecture/togaf9-doc/arch/chap...
I am saying when considering any “large” implementation there are a lot of considerations outside of infrastructure bills.
I’m not saying that every company should go cloud. But the “lenses” you have to look through are multifaceted
The bosses would’ve blamed me for choosing a tier 2/3 noname provider the first time a day of downtime happens. And they would’ve been right.
The better your competing offer, the better the negotiating position. And while Amazon hasn’t gotten more expensive per se, it’s certainly not gotten cheaper.
https://www.uber.com/en-GB/blog/up-portable-microservices-re...
My peak “wtf” moment was when we had a SEV because two services that should communicate actually used different versions of thrift, both hard forked by Uber, with different implementations for sets. Passing a set from one service to another caused everything to break.
I'd like to hear more about how Uber organized the engineering teams over two years to make "stateless microservices portable".
How many teams? What were the requirements to each team? What was the timeline? How did they know it was completed? How was it prioritized along other business priorities of the teams? How long did they think it would take originally? Was it worth it?
I’ve seen many teams go for simple/leaky abstractions on top Kubernetes to provide a similar solution, which is tempting because it’s easy and flexible. The problem is then all your devs need to be trained in all the complexities of Kubernetes deployments anyway. Hopefully Uber abstracted away Kubernetes and Mesos enough to be worthwhile, and they have a great infra team to support the devs.
They all use GitOps which means all infra deployments and changes are tracked and easily able to be rolled back on any issues. And the complexity is nothing compared to having to manage your own cloud resources using Terraform etc which used to be the case.
And these days every developer needs to be on board with DevOps and so there are no real old-school infra teams supporting anyone.
This is a major problem for databases and ultimately makes database "portability"/fault tolerance tricky since they work best with direct-attach storage that's inherently bound to a single physical machine.
I don’t know if we can truly abstract away the underlying system. The best we can do is give a best effort approximation that works in most cases, but explicitly call out the limits when they are reached.
I suspect that this is just the bubbling up of the underlying physics limitation of having limited resources where compute is run.
The implicit goal of these abstractions is really to central knowledge and best practices around the underlying tech. Kubernetes itself is trying to free developers from understanding server management, but you could argue it’s not worth using directly vs. just teaching your devs how to manage VMs for the vast majority of organizations.
I don’t think you’re ever going to stop more and more layers of abstraction, so the best we can hope for is they’re done well. Otherwise you may as well go back to writing raw ethernet frames in assembly on bare metal.
I disagree that the solution is to simply build more. Often the best thing to do is accept that devs will need to know a little infra, and work with that assumption.
> The implicit goal of these abstractions is really to central knowledge and best practices around the underlying tech.
I agree with that.
> Kubernetes itself is trying to free developers from understanding server management, but you could argue it’s not worth using directly vs. just teaching your devs how to manage VMs for the vast majority of organizations.
The difference is that spinning up a VM and setting it up to have all the features you would want from k8s would be too much to ask from a dev. You would probably just end up re-creating k8s.
> I don’t think you’re ever going to stop more and more layers of abstraction, so the best we can hope for is they’re done well. Otherwise you may as well go back to writing raw ethernet frames in assembly on bare metal.
The problem is that abstractions are not free, and most of the time they aren't done well. Once in a while you'll get one that reduces(hides) complexity and becomes an industry standard, making it a no-brainer to adopt, but most of your in-house abstractions are just going to make your life worse.
e.g. with kubernetes, if you have the actual manifests defined by every team, it is a pain to do any sort of k8s updates. With a simple abstraction where teams only define the things they are interested in configuring (eg helm values), that simplifies this task a lot.
Because engineers don’t have to understand infra, it often spans geographies and failure domains in unanticipated, undetectable ways. In my opinion the only antidote is a thorough understanding of your stack down to the metal it’s running on.
Even in a 100 person startup that I worked for where I designed the infrastructure and the best practices and wrote the initial proof of concept code and best practices for about 15 microservices it got to the point where I couldn’t understand everything and had to hire people to separate out the responsibilities.
We sold access to micro services to large health care organizations for their websites and mobile app's. We aggregated publicly available data on providers like licenses, education etc.
Our scaling stood up as we added clients that could increase demand by 20% overnight and when a little worldwide pandemic happened in 2020 causing our traffic to spike
If you are developing applications software like Uber 99.99% of the time you really do not need to be doing anything “fancy” or “exotic” in your service. Your service receives data, does some stuff with it (connects to a db or issues calls to other services), returns data. If you let those 0.01% of the things dictate where your internal platform falls on that spectrum, you will make things much more complicated and difficult for 99.99% of the other stuff. Those are where leaky abstractions and bugs come from, both from the platform trying to be more general than it needs to be and from pushing poorly understand boilerplate tasks (like configuring auth, certifications, TLS manually for each service) to infrastructure users.
Being unaware (of course not completely unaware, but essentially not needing to actively consider it while doing things) of infrastructure is actually the ideal state, provided that lack of awareness is because “it just works so well it doesn’t need to be considered”. It means that it lets people get shit done without pushing configuration and leaky abstractions onto them.
I’ll give you one example of something that does an excellent job of this: Linux. Application memory in linux requires some very complex work under the hood, but it has decent default configurations with only a couple commonly changed parameters that most applications don’t need much, and it had a very simple API for applications to interface with. Similar with send/receive syscalls and the use of files for I/O ranging from remote networking to IPC to local disk. These are wonderful APIs and abstractions that simplify very hard problems. The problem with in-house abstraction isn’t that they are trying to do abstractions but that sometimes they just don’t do a good job or churn through them faster than it takes them to stabilize.
> abstractions on top Kubernetes
> abstracted away Kubernetes
I am beginning to think it's not such a bad thing to live and work in a third-world country far away from SV-induced hype cycles. This is genuinely painful to read.Not every software development problem is a big scale problem and once you identify such a case you can start optimization work taking all the low level details into account. In reality most scalability problems revolve around databases, caches, concurrency and locks and you probably aren't going to tackle a lot of these in your average stateless service.
We've had individual EC2 instances go bad where I currently work, with Amazon acknowledging a hardware problem after a ticket is raised. The reality is, quickly resolving the issue means detecting it and moving off of the physical machine.
Naturally our tooling has no convenient way to do that, because we have layers of things trying to pretend physical machines don't matter.
The error rate on the machines was higher in both cases, but many requests still succeeded. Amazon certainly didn't detect an issue right away either.
I’m assuming this isn’t a web server, if so it’s even simpler.
/s
If management values business SLAs that is.
The answer to that is probably yes. APIs let us split work across systems/people/teams/regions, and provide a way for both sides of a split to work together. Uber has a lot of teams, a lot of engineers, and so it makes sense that there are a lot of API boundaries to allow them to work together more efficiently. Sometimes those APIs make sense to package as microservices.
And the thing is that nobody ever needs to run the entire stack other than end to end tests which get run in the cloud.
You just checkout the services you need and because they are designed to be isolated the dependencies will usually be automatically stubbed out. So it's just a matter of running them or chaining them together if you have a particular scenario to test.
I like to view k8s a lot like erlang's OTP, if something isn't right with the state of a service, I advocate calling 'exit()' and letting the restart with exponential backoff handle the transient.
Q2: How do you ensure the stubbed deps behave like the real thing?
Q3: how do you handle logging and metrics in an unified way across the stack? And related to this: how do you ever get to upgrade services crosscutting concerns that ideally are not invented in every service?
SRE here. Generally speaking, each API or each service will have a contract that it must adhere to depending on upstream and downstream relationships and their fail safes. Each service (or API) will then load test in isolation.
After that, if you want to be really sure about regressions (which would include fail safes) you load test the whole thing put together.
> Is the key some sort of meta tooling that understands relationships between microservices?
This is quite hard to do when you have a lot of transactions. I don't think there's commodity software that does this because you'd need to configure that software to map on keys, then map those keys to services. Generally, the easiest way is to get engineering teams to declare upstreams and downstreams.
> Q2: How do you ensure the stubbed deps behave like the real thing?
Generally, generation. Something like protobuf or Open API generation will do.
> Q3: how do you handle logging and metrics in an unified way across the stack?
You issue high level standards like, "We'll use JSON logging with UTC time formatting". At the end of the day logging is very contextual and in a service ownership model the service owners are usually the ones reading and alerting on their logs.
> And related to this: how do you ever get to upgrade services crosscutting concerns that ideally are not invented in every service?
Shared dependencies. I'm not actually sure what's a cross cutting concern; generally services that are this small should be designed to operate mostly independently. They're small, but "microservices" tend to have a lot of fail safes built in. If you're referring to how do we not write 4000 config loaders then there's usually a team that builds a very generic config loader and everyone or a majority use it.
perhaps simplest, biggest impact in my log life has been adhering these principles.
Q1: (perf) these tools exist, the buzzword phrase is "distributed tracing". The relationships are actually not explicitly defined for the tooling to work, but rather inferred. Visualize a network call as a call-stack, where each service is a level in the stack. Jaeger (a CNCF project addressing distributed tracing) was coincidentally started by Uber.
Q2 (stubs): In my experience, mocked responses get you a long, long way. Typically the API response type that you're mocking is generated from a protobuf (or thrift, OpenAPI, etc.) file. If your dependency changes that type in a way that breaks your test, the CI platform will let them know.
If it's a more subtle change (like, it used to deterministically return 18 and now it deterministically returns 20), it's really on the service owners to communicate changes and grep the code base before making the change.
Q3 (logging/metrics): Typically by using shared "logging" and "metrics" lib for each language. Every service will typically be a gRPC service and accordingly a standardized + generated-from-protobufs set of metrics to Prometheus, by default.
Q4 (how to upgrade common libraries): this is definitely a tricky one. The answer is, basically, really carefully. Typically, you'll want your infrastructure to be compatible with vX and vX+1, and give teams a deadline to cut over from logging X to X+1. The couple of weeks before that deadline usually involves a lot of cat-herding and handwringing.
But you tend to try to write your service so that it treats everything else it depends on like a vendor-provided API. Like, if you were building a Slack bot, you wouldn't ask Slack to let you pull down and run a local copy of Slack's API to test against. You'd maybe set up a test account in Slack's production system, and run your local bot against that to test it before you deploy it with credentials to run against your real slack account.
In a microservice architecture, you integrate with other internal systems in the same way.
I find the opaqueness of other services to reduce development speed quite drastically. With local code I can view both sides of the fence and easily see if I'm using it wrong or if it's a bug in my colleagues code.
Seems that if you're constantly developing against opaque services you'd end up in the same quagmire quite quickly?
But when things don't work as I expect, it's far more efficient to be able to view the code on both sides, rather than only on my own side and try to guess what the other side is doing.
Besides the usual suspect of wrong understanding on my end leading to misuse, this can also be due to lacking or wrong documentation of the other system, or bugs in the other system due to unexpected inputs or similar.
Like just a few days ago we spent an unreasonable amount of time with an API of one of our customers, where we would get empty list back for some of our queries. Turned out something in their service crashed when handed national characters, despite accepting JSON and hence UTF-8 input and nothing in the documentation about English letters only. Rather than returning 400 or 500, the service returned 200 with an empty list, leading us to assume we did something wrong.
Are you able to view the source code of your platform vendor? Everyone is at some level dependent on Black box APIs.
If you can document where with certain input you don’t get the expected output, you reach out to the team that is responsible for it whether internally or externally and they either explain it or they fix it.
This is the API service I’ve been working with over the past five+ years - three actually working at AWS (Professional Services).
https://boto3.amazonaws.com/v1/documentation/api/latest/inde...
I found a bug in one relatively new API that a service team released, I reached out to the team with a documented scenario and they fixed it.
Other times they explained what I was doing wrong. That’s what any large organization does.
I’ve worked with other vendors and internal teams plenty of times over the years.
Uber employees’ apps are special and allow us to log in as these fake users and create fake rides or deliveries, and then we can look at the traces and logs to debug and stuff.
Isn’t that the premise of the question? Does Uber need so many engineers?
The only people who can answer that are employees at Uber.
Don't even get me started on anything money related :)
I was with you on other types of users, but can you elaborate on these particular use cases?
"What I Wish I Had Known Before Scaling Uber to 1000 Services" - https://youtu.be/kb-m2fasdDY
My experience with microservice shops is you have one macromonolith with 50 people working on it (which has all the problems of a monolith and none of the benefits), 5 actual decent microservices with a team or individual that properly maintains them, and 100 random utility micro"services" that are like 3 lines of code, used by exactly one other service, and you need 40 loc and a network call to interact with them.
I'll take everyone has their own service any day of the week. At least when I need to interface with 12 different things I can have 12 different people to roast for not properly documenting their API. And tbh literally the only positive I can come up with for microservices is the ability to neatly fire one into the sun and rewrite it from scratch.
If all their It needs are behind micro "micro" services, that figure is understandable.
Outside of the map, taxi, food, payments, onboarding, they also have monitoring, deployment, HR, billing, legal, taxes, internationalized stufd, and the usual "..." for what I'm missing.
If you just take a standard ERP, you could easily split it in dozens even hundreds of microservices.
I call them nano services.
Everybody knows what is a monolith but nobody really knows what is the size of a "micro" service.
Just for taxes, do you make one service for taxes or one for each recipient of taxes? (In the EU, is it one for each country, in US, one for each state + federal ) with a different team managing each service?
So almost certainly they are duplicating their entire stack per-country if only to get around the vastly different regulatory environments.
But that scale introduces a lot of complexity so you can't just have "one service for onboarding drivers"
I find this hard to believe given the regulations from some of the larger countries requiring, by law, customer data be processed in country.
Responding comment says no they are not and the services are built to handle global traffic.
I respond and say I doubt that due to on soil laws.
You can argue two regions with different configurations but the same code bases are different services but that’s not what we’re talking about here.
Do you mean that your original point was about deployments to begin with?
FWIW I work in a microservices shop for a global app in an extremely regulation heavy industry, and we run a single codebase per service, segregating regulator-imposed behaviour via flags to deployments
Your fwiw is exactly what we’re taking about here and I’d venture a guess nearly half this site works for some Corp with duck tape, hope, and micro services powering their junk. Me too!
The same microservice that deployed in multiple geos still counts as one service, so considered to be 1 out of 4000 in this case.
That is a much easier business model and a lot lower level of complexity than Netflix. I imagine running pornhub is essentially running a large website that hosts video. Probably just the billing side of Netflix is more complicated than the entirety of Pornhub operation.
Uber's problem space is significantly more complex than Netflix, so I'm unsure it's a fair comparison. But they do seem to have quite a lot of overengineering going on. At least that's how I feel each time I read an Uber tech article.
About the only companies which seem to justify their complex architectures are Google/Meta/Amazon imo.
What makes you say this? Netflix serves probably several orders of magnitude more bytes and online video is hard. At its core Uber is basically a Passenger Service System and we had systems like these implemented in software since 1950s
Maybe a couple of dozens will be actual more complex and meaningful services. Then few dozens more services that are somewhat more unique.
And then majority of the long tail will be mostly cookie cutter services, doing X, but for lots of different use cases, where each of use cases is separate deployment counting as a service (for example - systems to process streams of logs related to business logic).
What would a monolith buy you?
An ability to easily change the boundaries of your conceptual components, because they WILL be wrong now or in the future.
Even with a well constructed monolith, you need to have well defined “services” with contractual interfaces.
You don’t have to understand 4000 services to make one change anymore than I need to understand the entire boto3 library when I am building on top of it.
https://boto3.amazonaws.com/v1/documentation/api/latest/inde...
You can’t just change your interface in a monolith either without breaking other parts of it.
The question is what makes you think 1 service is immediately better than however many payment services there are now?
Even Microsoft managed to do so for multiple products while also stack ranking the teams.
And I doubt there's a single service, even payments, that's as technically complex as Excel.
And I'd agree with the other child comment that the monolith can always be broken into separate components which are owned by different teams.
This sounds like a figure from someone who sees a signle microservice running across 100 pods/instances, and counted that as 100 "microservices".
We had to sign something at AWS not to divulge internal tooling like what we used for our internal account factory that we used to create AWS accounts. Literally tens of thousands of people know what this tool is.
It in fact was public.
https://aws.amazon.com/blogs/storage/how-automated-reasoning....
I verified that before I posted that little tidbit.
Orchestrating the application layer across clouds is interesting, but how does their data layer work?
I got so excited about reading for Mesos helping in the multi cloud world, potentially as the hypervisor for running k8s
But the underlying technology which carried them to this point is a fascinating read.
IMO, engineering man hour savings are a lot less trustable. This may eliminate or simplify some engineering processes but IME massive migrations like this simply replace them with a different set of processes; because they’re different and theoretically addressable they’re not counted against the hours saved as they can be bucketed into bugs/to be addressed by the roadmap/legacy behavior migrated from the old system (which is now dangerously-fragile-legacy and not ol-reliable-legac). Eventually someone will come along and decide this too is an inherently flawed platform that needs to be entirely replaced at great expense, and the circle of life continues.
This is still a massive undertaking not just from an engineering perspective but from an organizational/process one though. Whoever pulled this off essentially had to coordinate (or figure out how to simplify/explain things well enough to skip coordination) with almost every engineer and likely almost every production service in a company with thousands of engineers. Those in startups may balk about this kind of thing taking two years, but having done my own two year projects (at a smaller but comparable scale) in a big company I can say two years is what I’d consider a highly optimistic and unlikely outcome for a project of this magnitude.
Yes
> because they’re different
Now I have to learn an entire new set of tools/processes etc that are more useful to someone else but not helpful for me. The old one had its quirks but I knew it inside out and now the whole org has to re-learn how to do everything we did before.
Not defending their tech stack, but I mean that is a lot of realtime data that needs to be accurate - this is not your typical SaaS crud app.
And google is just a search engine they only need like 20 engineers……………
Is this generally a sign of youthful wishful thinking or just plain hubris?
We're giving a talk about this at KCD Denmark on the 14th of November "Keynote: Uber - Migrating 2 million CPU cores to Kubernetes" if anyone is in the area and has any particular interest in this.
I'm currently trying to get out of the industry because I'm drowning in architectural bullshit like this constantly. It is pedalled by snakes, bastards and wankers who care nothing for solving problems but want to create new ones.
Source: juggle lots of clusters full of things that shouldn't be in Kubernetes.
You can be all angry about it, but being angry at the storm doesn’t affect the storm, it only affects you. A lot. Negatively.
The trick is to position yourself to maximally profit from the next trend swing, I’ve been doing it for 20 years now, if you can predict where the next place is gold will fall from the sky, then go and stand there, with a really big bucket.
I always found it strange that there is a certain type of intellect who is capable of accurate observation of reality, but incapable of execution (sometimes called “the disconnected intellect”), they can tell you exactly the problem, and the solution, but sit angry/frustrated that the observed world doesn’t match some imagined ideal in their head, and rather than adapt their internal model and be entrepreneurial enough to capture the value that generates, they bleet and complain while losing all opportunities - opportunities they can see! I can’t imagine being like this.
I am fed up of solving the same problems again and again. It's more than just earning; there's intellectual dishonesty in this and it's tiring and demotivating.
I'm literally 18 months from packing up this shit and doing what I really want to do which involves nothing whatsoever to do with computers.
Where the pendulum will be in 2030? Asking for a friend :D
Put everything on: cost savings, energy reduction, privacy, death of advertising.
Cost savings -> inefficient languages and architectures will die because the main datacentre currency is going to be performance/watt and that's going to cost serious money when transport infra is contending with DC power consumption. Things which are compiled and not interpreted will have a cost benefit then. Rust/C# (with AOT)/Go etc. Half these bloated piles of shit with expensive build toolchains will die too.
Energy reduction -> linked with above, energy usage reduction is going to be a big one. That means reducing workforce, simplification and efficiency are going to be key drivers. This may kill some ML approaches off that consume a lot of energy. So ARM etc.
Privacy -> Confidence in surveillance states and the cloud is declining so privacy first oriented services are going to have a huge uptick. Apple / standalone systems / new opportunities.
Death of advertising -> advertising is in the death throes with AI coming in as it decreased the signal-to-noise ratio. It becomes less effective so discovery rather than promotion will be the way to get attention for your product. Portals / landing pages / software catalogues.
Me I'd concentrate on cloud cost management and code efficiency and business efficiency as key areas to invest my time in.
Bad faith is easiest seen when they hold opposition to higher standards than they can hit. In particular, on purpose. It is a smoke and mirror dialogue that is not intended to make progress, but only to spend the other side's time.
In contrast, being wrong and or not fully grasping difficulties is normal. And sometimes, you get lucky and unexpectedly make progress
Also, they are assuming their participation is a net positive, which is a position that requires some amount of intention to take, so I disagree that intention as you've cast it is really a relevant perspective here. Bad faith is more than just a rhetorical debate tactic, it's a modus operandi with regard to how someone actually engages with the topic at hand.
There is also a bit of the established problems obfuscating themselves. Such that it is easy to see many new workers have been given the run around many times.
2013: "Migrating Uber from MySQL to PostgreSQL"[1]
2016: "Why Uber Engineering Switched from Postgres to MySQL"[2]
[1] https://www.yumpu.com/en/document/view/53683323/migrating-ub...
[2] https://www.uber.com/en-GB/blog/postgres-to-mysql-migration/