The answer to that is probably yes. APIs let us split work across systems/people/teams/regions, and provide a way for both sides of a split to work together. Uber has a lot of teams, a lot of engineers, and so it makes sense that there are a lot of API boundaries to allow them to work together more efficiently. Sometimes those APIs make sense to package as microservices.
And the thing is that nobody ever needs to run the entire stack other than end to end tests which get run in the cloud.
You just checkout the services you need and because they are designed to be isolated the dependencies will usually be automatically stubbed out. So it's just a matter of running them or chaining them together if you have a particular scenario to test.
I like to view k8s a lot like erlang's OTP, if something isn't right with the state of a service, I advocate calling 'exit()' and letting the restart with exponential backoff handle the transient.
Q2: How do you ensure the stubbed deps behave like the real thing?
Q3: how do you handle logging and metrics in an unified way across the stack? And related to this: how do you ever get to upgrade services crosscutting concerns that ideally are not invented in every service?
SRE here. Generally speaking, each API or each service will have a contract that it must adhere to depending on upstream and downstream relationships and their fail safes. Each service (or API) will then load test in isolation.
After that, if you want to be really sure about regressions (which would include fail safes) you load test the whole thing put together.
> Is the key some sort of meta tooling that understands relationships between microservices?
This is quite hard to do when you have a lot of transactions. I don't think there's commodity software that does this because you'd need to configure that software to map on keys, then map those keys to services. Generally, the easiest way is to get engineering teams to declare upstreams and downstreams.
> Q2: How do you ensure the stubbed deps behave like the real thing?
Generally, generation. Something like protobuf or Open API generation will do.
> Q3: how do you handle logging and metrics in an unified way across the stack?
You issue high level standards like, "We'll use JSON logging with UTC time formatting". At the end of the day logging is very contextual and in a service ownership model the service owners are usually the ones reading and alerting on their logs.
> And related to this: how do you ever get to upgrade services crosscutting concerns that ideally are not invented in every service?
Shared dependencies. I'm not actually sure what's a cross cutting concern; generally services that are this small should be designed to operate mostly independently. They're small, but "microservices" tend to have a lot of fail safes built in. If you're referring to how do we not write 4000 config loaders then there's usually a team that builds a very generic config loader and everyone or a majority use it.
perhaps simplest, biggest impact in my log life has been adhering these principles.
Q1: (perf) these tools exist, the buzzword phrase is "distributed tracing". The relationships are actually not explicitly defined for the tooling to work, but rather inferred. Visualize a network call as a call-stack, where each service is a level in the stack. Jaeger (a CNCF project addressing distributed tracing) was coincidentally started by Uber.
Q2 (stubs): In my experience, mocked responses get you a long, long way. Typically the API response type that you're mocking is generated from a protobuf (or thrift, OpenAPI, etc.) file. If your dependency changes that type in a way that breaks your test, the CI platform will let them know.
If it's a more subtle change (like, it used to deterministically return 18 and now it deterministically returns 20), it's really on the service owners to communicate changes and grep the code base before making the change.
Q3 (logging/metrics): Typically by using shared "logging" and "metrics" lib for each language. Every service will typically be a gRPC service and accordingly a standardized + generated-from-protobufs set of metrics to Prometheus, by default.
Q4 (how to upgrade common libraries): this is definitely a tricky one. The answer is, basically, really carefully. Typically, you'll want your infrastructure to be compatible with vX and vX+1, and give teams a deadline to cut over from logging X to X+1. The couple of weeks before that deadline usually involves a lot of cat-herding and handwringing.
But you tend to try to write your service so that it treats everything else it depends on like a vendor-provided API. Like, if you were building a Slack bot, you wouldn't ask Slack to let you pull down and run a local copy of Slack's API to test against. You'd maybe set up a test account in Slack's production system, and run your local bot against that to test it before you deploy it with credentials to run against your real slack account.
In a microservice architecture, you integrate with other internal systems in the same way.
I find the opaqueness of other services to reduce development speed quite drastically. With local code I can view both sides of the fence and easily see if I'm using it wrong or if it's a bug in my colleagues code.
Seems that if you're constantly developing against opaque services you'd end up in the same quagmire quite quickly?
But when things don't work as I expect, it's far more efficient to be able to view the code on both sides, rather than only on my own side and try to guess what the other side is doing.
Besides the usual suspect of wrong understanding on my end leading to misuse, this can also be due to lacking or wrong documentation of the other system, or bugs in the other system due to unexpected inputs or similar.
Like just a few days ago we spent an unreasonable amount of time with an API of one of our customers, where we would get empty list back for some of our queries. Turned out something in their service crashed when handed national characters, despite accepting JSON and hence UTF-8 input and nothing in the documentation about English letters only. Rather than returning 400 or 500, the service returned 200 with an empty list, leading us to assume we did something wrong.
Are you able to view the source code of your platform vendor? Everyone is at some level dependent on Black box APIs.
If you can document where with certain input you don’t get the expected output, you reach out to the team that is responsible for it whether internally or externally and they either explain it or they fix it.
This is the API service I’ve been working with over the past five+ years - three actually working at AWS (Professional Services).
https://boto3.amazonaws.com/v1/documentation/api/latest/inde...
I found a bug in one relatively new API that a service team released, I reached out to the team with a documented scenario and they fixed it.
Other times they explained what I was doing wrong. That’s what any large organization does.
I’ve worked with other vendors and internal teams plenty of times over the years.
Uber employees’ apps are special and allow us to log in as these fake users and create fake rides or deliveries, and then we can look at the traces and logs to debug and stuff.
Isn’t that the premise of the question? Does Uber need so many engineers?
The only people who can answer that are employees at Uber.
Don't even get me started on anything money related :)
I was with you on other types of users, but can you elaborate on these particular use cases?
What would a monolith buy you?
An ability to easily change the boundaries of your conceptual components, because they WILL be wrong now or in the future.
Even with a well constructed monolith, you need to have well defined “services” with contractual interfaces.
You don’t have to understand 4000 services to make one change anymore than I need to understand the entire boto3 library when I am building on top of it.
https://boto3.amazonaws.com/v1/documentation/api/latest/inde...
You can’t just change your interface in a monolith either without breaking other parts of it.
The question is what makes you think 1 service is immediately better than however many payment services there are now?
Even Microsoft managed to do so for multiple products while also stack ranking the teams.
And I doubt there's a single service, even payments, that's as technically complex as Excel.
And I'd agree with the other child comment that the monolith can always be broken into separate components which are owned by different teams.
So almost certainly they are duplicating their entire stack per-country if only to get around the vastly different regulatory environments.
But that scale introduces a lot of complexity so you can't just have "one service for onboarding drivers"
I find this hard to believe given the regulations from some of the larger countries requiring, by law, customer data be processed in country.
Responding comment says no they are not and the services are built to handle global traffic.
I respond and say I doubt that due to on soil laws.
You can argue two regions with different configurations but the same code bases are different services but that’s not what we’re talking about here.
Do you mean that your original point was about deployments to begin with?
FWIW I work in a microservices shop for a global app in an extremely regulation heavy industry, and we run a single codebase per service, segregating regulator-imposed behaviour via flags to deployments
Your fwiw is exactly what we’re taking about here and I’d venture a guess nearly half this site works for some Corp with duck tape, hope, and micro services powering their junk. Me too!
The same microservice that deployed in multiple geos still counts as one service, so considered to be 1 out of 4000 in this case.
Maybe a couple of dozens will be actual more complex and meaningful services. Then few dozens more services that are somewhat more unique.
And then majority of the long tail will be mostly cookie cutter services, doing X, but for lots of different use cases, where each of use cases is separate deployment counting as a service (for example - systems to process streams of logs related to business logic).
If all their It needs are behind micro "micro" services, that figure is understandable.
Outside of the map, taxi, food, payments, onboarding, they also have monitoring, deployment, HR, billing, legal, taxes, internationalized stufd, and the usual "..." for what I'm missing.
If you just take a standard ERP, you could easily split it in dozens even hundreds of microservices.
I call them nano services.
Everybody knows what is a monolith but nobody really knows what is the size of a "micro" service.
Just for taxes, do you make one service for taxes or one for each recipient of taxes? (In the EU, is it one for each country, in US, one for each state + federal ) with a different team managing each service?
"What I Wish I Had Known Before Scaling Uber to 1000 Services" - https://youtu.be/kb-m2fasdDY
That is a much easier business model and a lot lower level of complexity than Netflix. I imagine running pornhub is essentially running a large website that hosts video. Probably just the billing side of Netflix is more complicated than the entirety of Pornhub operation.
Uber's problem space is significantly more complex than Netflix, so I'm unsure it's a fair comparison. But they do seem to have quite a lot of overengineering going on. At least that's how I feel each time I read an Uber tech article.
About the only companies which seem to justify their complex architectures are Google/Meta/Amazon imo.
What makes you say this? Netflix serves probably several orders of magnitude more bytes and online video is hard. At its core Uber is basically a Passenger Service System and we had systems like these implemented in software since 1950s
My experience with microservice shops is you have one macromonolith with 50 people working on it (which has all the problems of a monolith and none of the benefits), 5 actual decent microservices with a team or individual that properly maintains them, and 100 random utility micro"services" that are like 3 lines of code, used by exactly one other service, and you need 40 loc and a network call to interact with them.
I'll take everyone has their own service any day of the week. At least when I need to interface with 12 different things I can have 12 different people to roast for not properly documenting their API. And tbh literally the only positive I can come up with for microservices is the ability to neatly fire one into the sun and rewrite it from scratch.
/s
If management values business SLAs that is.
This sounds like a figure from someone who sees a signle microservice running across 100 pods/instances, and counted that as 100 "microservices".
We had to sign something at AWS not to divulge internal tooling like what we used for our internal account factory that we used to create AWS accounts. Literally tens of thousands of people know what this tool is.
It in fact was public.
https://aws.amazon.com/blogs/storage/how-automated-reasoning....
I verified that before I posted that little tidbit.