Domain-Oriented Microservice Architecture
eng.uber.com
eng.uber.com
...no, not at all. There are a couple of operational benefits, but vastly more drawbacks, and on balance microservices are phenomenally harder to operate than monoliths.
Organizations adopt microservices when the logistical overhead of coordinating teams against a monolith becomes so large that it starts affecting product velocity. Microservices are a way to release that organizational/logistical friction, at great technical cost.
With that said, the domain-oriented bounded context is indeed the right way to think about service delineation.
We went microservices for the canonical reason you should - to be able to release and scale independently.
We were used by health care networks. Can you imagine the increased use post-Covid?
We even had some that were both hosted on Fargate (Serverless Docker) with lower latency, more expensive but slower scaling for online use and hosted on Lambda for internal batch use - higher latency, faster scaling, less expensive for batch use. The CI/CD pipeline deployed to both.
That instead of being able to granular take part of the application that had a larger load, and run it on a Firecracker micro VM (the underlying VM for lambda and Fargate) with 256MB RAM and 1 core, we add enough VMs to scale a monolithic app - even the parts that don’t need it on a full scale VM with 8GB RAM and four cores?
We did actually have to do something similar for a legacy Windows app. We scaled the entire process up based on the number of messages in the queue. It was extremely wasteful. It required at least a 4GB/2 CPU VM compared.
edit: If breaking out one service from your monolith worked for your use case, that's great. I'm not trying to deny your experience. It is atypical, however.
because a read replica actually is the same database application (same "monolith") just started with different parameters, to behave as a replica
It’s not just a matter of spinning up a database.
Also, since many enterprise apps live stores procedures and putting business logic in the database, that’s another ball of wax you have to untangle.
That's certainly one of the operational limits, but arguably not the most important one.
You distribute your system so that parts of it may scale independently, for example. Traffic fluctuates along time and the only option you have to scale your system is adding nodes as you go.
Microservices are a solution to organizational problems, not technical problems. On balance they create far more technical issues than they solve.
I've spent a lot of time trying to understand when a microservices architecture makes sense, what the caveats are, and what philosophy one should take to building services. All the material I've read seems to point in the direction of services being ideally coupled to domain boundaries.
It seems to me that Uber's services proliferated beyond the framing of bounded contexts, and DOMA is their attempt to reign it back in again. I think it's an excellent strategy, and arguably a very good approach for other companies who find themselves in this position.
I don't think DOMA is a good place to stay at. The network should only be tolerated as long as it provides benefits that outweigh the costs. "Monolith" is not synonymous with "poor design". Seeing that these enclaves of services sitting within a domain depend on eachother in the way the OP describes, it really makes me think that they'd find further benefits by expelling the network from each domain.
I agree. There should be no need to use network calls to enforce interface boundaries if you have a cohesive bounded context.
Or to be snarky about it, welcome to 2004 Uber! Eric Evans sends his regards!
There will always be people who don’t want to respect where the boundaries are drawn (not in a constructive, “it could be better” kind of way, but in a “but this works, too” kind of way). If a group of such people get together, microservices compartmentalizes their capacity to drag down the ship, so to speak. I think this risk compartmentalization is a benefit that must be weighed against the costs (in terms of latency, maintenance of shared libs, opentracing, etc). These days the costs are vanishing, as tools are quite good and becoming easier to manage.
All that said, if you’re a small team of senior engineers working with a shared mental model, a single binary with internally-bounded contexts works really well and I agree with you, having seen it done well.
Fortunately everyone bought into the architecture, and respected the boundaries. Not everyone was senior, and not all the code was great, but we adopted the viewpoint that so long as the bad code is in the right spot (and not talking to things it shouldn’t) everything would be ok in the end. And it was.
Half a dozen teams working in one codebase was definitely pushing it though, and the need to scale much beyond that would have definitely required some service-level compartmentalization to keep the ship from sinking, as you said.
Even then, it would still be a far cry from the “microservices should be small enough to re-write in 2 weeks” approach.
I’ve seen many times where people were afraid to change boundaries because they assumed the first person got the architecture exactly right.
There aren’t any major drawbacks to this model when a business is young (first couple years). The downsides appear when you have different parts of your application with very different load requirements.
It also takes a lot of discipline to write code this way. Without strict code review and more experienced hands, the bounded contexts fall apart.
One of the advantages of the microservices model is it limits the damage people can do. :)
It forces a bounded context on a team of engineers and says “hey, play in this sandbox and follow these SLAs. If your internal designs are awful, good luck.”
So to me the OP reads like "we're coming up with some new terminology for a bounded context, and also defining how those contexts should be allowed to layer in order to simplify/control failure modes".
The layering stuff is more interesting than Uber's rediscovering bounded contexts, though it's definitely interesting that they have come into agreement with Evans (and the rest of the DDD community) on the "Service == Bounded Context" principle.
You're quite right.
> [...] defining how those contexts should be allowed to layer [...]
This, however, seems to go against the spirit of things. There is a consistent "ubiquitous language" within a bounded context, where domain terms are concrete and unambiguous. (Or rather, the context disambiguates the language.) The concept of "layered contexts" seems to neuter the concept. Does each layer successively disambiguate the one above? Or does it add new terms that didn't exist?
The layering here sounds much more technically-motivated than domain-oriented. And my argument is that the networking internal to their "domains" is largely an artifact of having build DOMA out of a plethora of disorganized microservices. Doubtless there will be some necessary networking remaining, as you remarked on, such as between processors and databases. But the origin here suggests most of it is left over from what came before.
I'm not sure how many systems have been built using DDD with hundreds of interacting bounded contexts, but I suppose I could believe that _some_ structure would be beneficial. (If you know of any case studies here I'd love to hear of them, I've not actually seen anything published on this topic.)
In general the concept of an "infrastructure bounded context" seemed a bit weird to me from my understanding of DDD, but then I though about Kubernetes, and you could make a case that it is an example of such a bounded context; it has its own ubiquitous language, etc. It would be weird for your infrastructure to have any understanding of the domain objects running on top of it, so a hierarchy makes sense.
Likewise if you have BFFs for your different API clients; the domain services underneath them could be abstracted away from things like REST, if all your internal services use gRPC (for example). You could consider this the UI layer in DDD's layered architecture.
I'm struggling to come up with more sensical layers than that though; in DDD there's the Application and Domain layers; I don't really see how you'd pull "Application" vs. "Domain" bounded context layers together in a way that made sense.
> But the origin here suggests most of it is left over from what came before.
I'd certainly agree with this -- it seems like lots of the intra-BC complexity is excessive compared to what you'd get if you built your services with a BC in mind from the start.
I don't think I'd emulate their intra-BC structure, it's only the inter-BC organization that I think has any merit for other systems (and even then I'm not fully convinced yet).
I know I'm being snarky and I have used micro-services myself, but only when it was smacking me in the face as the best tool for the job. Is the fools-gold rush still on to do everything as a micro-service from the get-go?
Side note, I can't wait to hear my boss use "DOMA" in a meeting in the coming months. FML.
Yes.
For global-scale web applications? Obviously you won't. High-availability, low latency, resilience, scalability, performance. You don't get any of that by running your app on a single box. That ship has sailed two or three decades ago. Physics establishes all the limits, not software architects.
Distributed system critics, where they fixate on trendy microservices architectures or poopoo other suggestions like DOMA, should take a step back and look at themselves and what they are actually complaining about. Yes, a solution with no moving parts is simpler than a solution with some moving parts. But have you really noticed what problems are being solved by adding these pieces?
They are all, however, running exactly the same code, just in different configuration. I'd say it's roughly speaking 80% core libraries and the remaining 20% varies by role.
I have a command-line control & observation tool that comes in bin/ of the same repository, and it is again wrapped around the same code besides.
That's the modern monolith in production.
I mean, you yourself talk about "different instance roles".
You should pay attention to your own claims: if you have a distributed deployment comprised of different nodes, and you have specialized nodes that you yourself state that are ran to handle limited and very specific responsibilities, then just because you decided, for any reason that only you can think about, to bundle everything in a single project... That doesn't make it a monolith, does it?
Just to be absolutely clear, "monolith" is not a reference to how you chose to organize your source tree. Monolith is a software architecture concept that defines how your whole application is organized and deployed. A distributed system comprised of multiple specialized processeses running independently is not a monolith, even if you somehow believed it was a good idea to pick which role you run through configuration.
> you yourself
> You should pay attention
> you decided
> you yourself state
> even if you somehow believed it was a good idea
This isn't language anyone should respond to, and not merely because it's staking out a fine example of the No True Scotsman fallacy.
> you yourself state that are ran to handle limited and very specific responsibilities
I didn't.
> Just to be absolutely clear, "monolith" is not a reference to how you chose to organize your source tree
Again, that's a straw man - I never said it was. Although I can certainly see how someone who was absolutely determined to make an unnecessarily bitter remonstration as personal as possible might - through either branch of Hanlon's razor - misconstrue the words "the same repository" adversarially for the purposes of their ego trip.
It's a monolithic application because any of the instances could perform any of the roles, and they're all running exactly the same code. They're distinguished in production for the purposes of operational sanity, because only a flaming idiot would, say, run reporting workloads on the API host.
But when I stand up a demo / showcase environment, for example, it has exactly one instance that does everything, and we can (and do) develop with the whole thing running single process on our laptops.
I shouldn't need to clarify any of this, because the point being made was a rebuttal to the "single box" thesis, not whether I met some gatekeeper's opinion about my standing to discuss the topic.
The empirical example is that, say, Shopify and Basecamp both described their applications as monolithic.
However, there's a more comprehensive demonstration, in that we can't define "box" without contradiction. When you consider the many shells of virtualisation we use, the fact that any web application is by definition accessed over a network (and therefore distributed), and the internal architecture of a modern server, which is practically a distributed federation, or even the coordination between threads in a single process, it leads inevitably to contradiction (if you're careful) or just a messy quagmire of conflicting definitions.
The final nail in this dichotomous coffin is that the converse is also untrue, since any microservice-based application can be deployed on one box. Whichever way you look through the scope, deployment model turns out orthogonal to the taxonomy of software architecture.
Quite obviously the point is not whether a microservice-based app can be deployed on a single box. It obviously can.
The point is that you split your app in order to deploy different pieces to different boxes then you have a distributed system and no longer a monolith.
A monolith approach means a single program. This is a software architecture approach and implies by definition that the whole app runs as one in a single box.
I don't understand the point of arguing.
There is indeed no point in arguing about that.
And I’m not even sure about the calculator.
List of stacks I have used in some form since then, raw TCP/IP for in-house RPC protocol, SUN RPC, RMI, COM/DCOM, XML-RPC, SOAP, CORBA, REST, WCF and apparently gRPC is the new fashion.
At the same time, I also done modular development with teams responsible for modules, where the language features for creating modules, defining interfaces, and use binary dependencies are taken into use.
What I usually see with most "distributed systems" is that they are used as a physical solution for teams that never written a proper module in their life.
If monoliths with total lack of modularity are hard to debug, spaghetti network calls are even less fun.
So yeah, we keep going at this, and then people discover that monoliths written in a correct modular way, with libraries, happen to be easier to debug and reason about without a network in the middle.
Main problem seems to be that not many developers bother to read about modular programming, large scale development (like Lakos books) and what features their language of choice offers for such endeavours.
"Large-Scale C++ Software Design"
https://www.amazon.com/Large-Scale-Software-Design-John-Lako...
Although oriented towards C++, many architecture tips apply to other languages as well.
John Lakos is in the process of writing updated versions of the book.
"Large-Scale C++ Volume I: Process and Architecture"
https://www.amazon.com/Large-Scale-Architecture-Addison-Wesl...
"Large-Scale C++ Volume II: Design and Implementation"
https://www.amazon.com/Large-Scale-Implementation-Addison-We...
Then going back into the old days, you have
"Software Engineering in Modula 2: An Object Oriented Approach"
https://www.amazon.de/-/en/Jill-Hewitt/dp/0333515188
"Data Structures and Program Design in Modula-2"
https://www.amazon.de/Larry-R-Nyhoff/dp/0023886218/ref=sr_1_...
"Code Complete: A Practical Handbook of Software Construction"
https://www.amazon.com/dp/0735619670/ref=sr_1_1
"AntiPatterns: Refactoring Software, Architectures, and Projects in Crisis"
https://www.amazon.com/AntiPatterns-William-J-Brown/dp/04711...
"Component Software: Beyond Object-Oriented Programming"
https://www.amazon.com/Component-Software-Object-Oriented-Pr...
"Use Cases Combined With Booch/Omt/Uml: Process and Products"
https://www.amazon.com/Use-Cases-Combined-Booch-Omt/dp/01372...
Just some pointers to get you started.
I have posted some links to books on sibling comment.
IMO Micro-Services were addressing the skill gap in designing comprehensive schemas, not so much the object layer between user and data. So not modular "programming" but rather "modular design".
Feels a bit like the Hexagonal Architecture being rediscovered.
Micro service with half life of 1.5 years? Does that mean that enough planning is not being done? Or leadership failure at software planning level?
Collaboration between teams at scale is very hard but that is what the leadership layer is for - to collaborate more not build more micro services.
This is literally scaled/distributed domain-driven design (DDD).
I have felt strongly that folks got so caught up in the hype of yet another new thing they forgot how to extend what came before.
It feels like a reinvention of existing ideas at larger scale, and a pattern we keep repeating.
I'm not complaining, of course - new things are possible and being learned through this innovation. I do feel we should be more careful on the cutting edge, to see how it relates to where we came from.
Put another way, reinventing the wheel is not pointless if you come out with a better thing, or better wheel. But don't forget what was good about the previous wheel before you throw it out?
I am making an observation - when each new "fad" or "hype cycle" tech starts, it seems as though the pattern knowledge of what came before is discarded, or, disregarded as "legacy" or possibly even just forgotten. It feels like a knowledge transfer is missing. It would be terrible if, we, as an industry aren't passing down knowledge and reinventing hard won pattern discoveries efficiently.
Did this pattern Uber discovered come from studying onion architecture, DDD, etc first, finding the limitations, and then scaling them?
Or did this arise from throwing away everything that came before (or not knowing about it), forging an undiscovered path, and then rediscovering the old patterns could be applied? If the latter, what can be learned to make this process of discovery and linkage to existing patterns more efficient?
I think this article shows innovation is tricky, or, the risk (and potential reward) at the bleeding edge. Leaving behind design constraints of what came before might be necessary.
Maybe I'm trying to say, as an industry we need to balance exploitation of previous knowledge and our attitudes about how we feel about "legacy", with the unquenchable thirst for the next new innovation?
The key to any service being usable outside of the exact context it was first written in is to ensure no product specific business logic is added to it.
By splitting out infrastructure from general business from product specific services, you can do a much better job of understanding, and therefore controlling where where product specific logic is allowed. This in turn will make your lower level services far less coupled to the exact context they are first used in.
For smaller orgs, this is by far the more useful information, rather than how to deal with 2200 microservices.
> Uber has grown to around 2,200 critical microservices
Even thinking about the very largest systems I’ve worked on, I can’t think of what could possibly be split into 2000 separate individual services. What are these thousands of microservices and how micro are they?
The amount of services is really quite meaninglessness compared to the amount of developers. As with anything the efficiency good down as the amount of people goes up. There is no way that 2000 people can agree on anything. So most likely there is lots and lots of overlap.
You could probably refactor some of this duplication out, but by the time you would be done new duplication would have emerged. Keeping many people in sync is just difficult :)
Anyway the point is that 2200 services is meaningless without telling how many employees are working on said services :)
I assure you there are /significantly/ more than 2200 instances of the services running.
They do indeed mean 2200 separate deployable services.
Not saying that's a good way to build systems but it's definitely one way.
Heck, simple auto insurance systems with multiple products in a single large state can hit 5k really easy.
Then you look at the verticals and there's maps, payments, analytics, hosting, marketing, security, identity, partnerships, third party integrations, and a million other things you don't think about.
It's pretty easy to get into those numbers if you want zero downtime and code your own verticals, and have the resources to do it.
For instance, in India you get autorickshaws (called "Autos" sometimes) in addition to cabs as a seperate option. Autos have different rules, billing, driver compliance and safety standards. There are probably a bunch of services around Autos.
Similarly, safety regulatons differ in each country, and in some cases teams are forced to act quickly. Having them in separate microservices again makes sense. Same goes for offers, eats etc.
2000 services is not inconceivable for a company operating across the world.
It’s worth pointing out that not all of the 2200 services may be user-facing. Some may be internal, such as admin tooling or CI services. That said, 2200 seems like a lot!
Here's Uber's announced results of $18 billion in gross bookings https://investor.uber.com/news-events/news/press-release-det...
Round that to significant figures and you get 10%.
I was employed by Uber (I quit after a couple months), and the idea's that they now hint towards were largely rejected by the engineers, and that was less than a year ago. Uber is just in the position to throw a large sum of money at making wrong decisions and getting away with it, because it's not their money. It's VC funny money.
DISCLAIMER: Yes, I did indeed create this account to be able to reply anonymously.
If you want to know how not to do things, Uber is a very good place to look.
Not that this is unique to Uber, mind. I just described half of my coworkers too.
People are always grabbing nouns and thinking they’re microservices.
"Previously product teams would have to call numerous downstream services to leverage a domain; they now have to call just one....Furthermore, we were able to classify 2200 microservices into 70 domains."
So they went from 2200 to 70 microservices? One extreme to the other. The answer is somewhere in the middle.
I think this is key sentence. Microservices does not make a lot of sense for teams of 10s, while are a great tool for team of 100s.
We are already at a point where you could hypothetically do this for a reasonably-sized organization. A 64 core CPU can support a huge number of clients. Stuff like .NET Core scales really well if you want to build something complex like this. One big binary that occupies an entire physical host is an extremely compelling development model. Literally everything becomes a direct method invocation. You can also have type enforcement and atomic releases for the entire enterprise. Also makes a monorepo an obvious choice for source code management.
It sounds like a complex monster until you build it one time. Then you are basically done. The leverage you get when you have 1 way to do everything is extremely powerful. I do recognize there are scenarios where you cant force one persistence abstraction on all use cases, but there's no reason you couldn't have a TimeSeriesEntity (keyed by time) in addition to a typical BusinessEntity (keyed by a unique integer). Both could have unique replication implementations, but it would still be standardized and all nodes would be speaking the same protocol because they all derive from the same source.
The monolith only remains a monolith in operational terms until the engineering team develops some imagination.
[1]: https://www.cs.ubc.ca/~gregor/teaching/papers/4+1view-archit...
I must be honest: aside from the context in which it was linked, I came away terribly unimpressed by this paper. I found myself disagreeing on a number of points. Having only spent half an hour on it, I can't claim to have a major problem with it, but I wouldn't couch my architectural work against it.
One thing that was not touched on - how is security done in DOMA?
* Is auth happening at domain gateways or at each service?
* Similarly for encryption, does it terminate at Domain GW or at each service?