Adopting Microservices at Netflix: Lessons for Architectural Design
nginx.com
nginx.com
> One kind of coupling that people tend to overlook as they transition to a microservices architecture is database coupling, where all services talk to the same database and updating a service means changing the schema. You need to split the database up and denormalize it.
That sounds like a decision you wouldn't want to take lightly; the kind of thing you might do once your company is already big. I wouldn't want to start out that way though, it sounds like a recipe for a mess.
Does anyone know?
Is the best practice to not use entity services?
For updates and maybe reads, it's a different story.
\d Person
[A bunch of table schema stuff]
Referenced by:
TABLE "Billing" CONSTRAINT "billing_id_fk" FOREIGN KEY (id) REFERENCES Person(id) ON DELETE CASCADE
(I typed that off the type of my head so it might not be quite correct.)
So the basic premise is that there is no shared database, and thus having the database enforce cascading deletes is not an option.
The Billing service can then reference People or Businesses (and if Businesses, then sub-People), bill to an Address, etc.
No one's saying every object should be a service; you need to find the correct lines to divide across.
In our system (which has been service-oriented for five years), we don't do deletes. We do 'inactive' (UPDATE table SET ACTIVE=0…), but never deletes.
Especially in a case of your billing example, you never want to delete a person or address, because that's historical data you need to retain, but we just keep everything. If it goes in the database, it's because we want to keep it forever.
You'd have to allow for propagation delay. Plus the possibility of a message storm if you delete something fairly fundamental.
You're still screwed if you complete a transaction on the deleted person's still-existing account, now that your system is no longer transactional...
It adds complexity of course but that is the give/take when you use microservices.
I think you basically have to learn to live in an eventually consistent world. In the case of people being deleted I would imagine that the user service exposes a pub/sub interface where address and billing services subscribe to "delete" events.
http://adrianmarriott.net/logosroot/papers/LifeBeyondTxns.pd...
The problem of services being up/down has been solved with service discovery e.g. Consul, Etcd, Zookeeper.
If System A needs to tell System B about an event in order for A and B to remain consistent, but B is down, you've got eventual consistency, because B can't become consistent with A until it's back up and has performed whatever recovery is necessary to process that event. Service discovery does nothing to solve that problem.
In addition, the network isn't just up or down. It's varying shades (dare I say, 50 shades?) of down or broken. A single machine might not be accessible due to a switch issue. An entire rack or aisle might be compromised by a bad router or faulty routing table. A network cable might be flaky. The truth is you just don't know, and that's all inside a single LAN.
Your service discovery system could be able to see service {A,B,C}, but service A can't talk to B or C due to network issues. It happens.
It explains how to manage situations like this.
One may argue that in this case, "people" service doesn't go by description of micro-service as given in the article. But we need to understand that services get called in some context and there has to be someone there to do the plumbing. That someone can either be a db query, some code in the service, or app/application calling the services. And "generally" you would prefer service code over other two and hence a composite service.
IMO it may also be okay to have People, address, billing under one schema if service granularity and context allows so.
It's tempting because it's easy, and at first glance it seems to solve lots of problems (consistency, communication etc).
It's a big mistake.
If multiple service are reading and writing from the same DB it rapidly becomes impossible to change things.
Things like input validation changes suddenly have to be implemented in multiple places (which is hard), and done simultaneously (which often becomes hard enough to stop work being done).
In a microservice based approach, the data flows though shared services, and so changes can be done in fewer places.
(Note that read-only, reporting-style databases are separate. I think there is a good case for these being shared)
You can create a tightly coupled design with a shared DB, but there is nothing inherent with shared DB integration requires that.
This has always been the generally accepted way to scale out software services. Is there a novel idea being discussed here, or just that they've been doing this at Netflix?
And there was a whole hubbub of discovery what with UDDI and DISCO and all that jazz I never really understood.
It'd be nice if they gave some solid examples of the "micro" part to distinguish it from the general SOA idea.
Cockcroft defines a microservices architecture as a service-oriented architecture composed of loosely coupled elements that have bounded contexts.
To me this just reads as "service-oriented architecture in a sensible way".
No one thinks to intentionally build a SOA with tightly coupled components with poor boundaries.
It's a great architecture, but fan out of dependencies is a real risk.
The load balancer will spin up more machines, so now you have 10 machines leaning on whatever the back end is.
Yes, your approach is great - but you really have to understand the failure modes - if you're living on the edge, you could have a pretty un fun cascading error.
The usual alternative is redeploying and scaling the entire app to iterate on performance, which is a much slower process. If performance is a concern, microservices should be a big win.
Bringing those dependencies in under one roof doesn't eliminate their risk. Though it does increase the chance that an error in one of these small dependencies brings down the whole system.
With microservices if one of the services runs into trouble, you can ignore it and still serve the other 90-99% of your site without it. You can also deploy updates to services without having to deploy your entire site.
Breaking your system out into multiple dependencies means you can not only scale your infrastructure, but you can scale individual parts of your infrastructure based on demand, bottlenecks, usage, etc.
Netflix has talked about in the past how, because their systems are broken apart, they don't have to deal with these issues. Rating service having problems? Don't show user ratings. Search service offline for updates? Disable search. If Netflix was one giant (Rails? Django? Node?) app, it would be very difficult to cut out poorly-performing parts temporarily.
As an example of a (probably?) bad way to organize services, I worked on a project that had factored a role-based access control system into its own service. Every single web request hit this service, which made it a single point of failure, performance critical, impossible to temporarily disable, etc.
If you weave together 20 services to produce one mega-service, it's much harder to optimize for performance and even just keep the implementation correct. Caching of complicated multi-factor answers is less frequently possible.
Also, 20 microservices may feed e.g. 5 large 'end-user' services together, in various combinations. If one of the microservices is slow, only these of the end-user services are affected that actually use it.
Monolithic large services are harder to combine, so the risk that an unrelated remote service somehow gets called in the process and slows things down is higher.
All these things are good. You want isolated, focused test environments. You want tightly defined alarms. However we underestimated how long creating a new service would take. In the end we ended up pushing features out when they were ready but before the operational work was complete. Unsurprisingly we saw the issues we knew we wanted to protect against.
Better microservice franeworks that match the companies infrastructure would be helpful. Make building microservices cheap by building tools to speed up the process.
I hope micro services is not just a new fashion in software and is actually useful ten years from now.
a) How do you prevent technical debt? It seems to be more difficult due to APIs which shouldn't have breaking changes. In theory you could always version up the APIs and serve both versions or just add a new API for a breaking change, but these solutions seems awkward.
b) How do you start developing multiple microservices at the same time? I would expect APIs to change a lot in the beginning, which would mean that updating one microservice would break another. Perhaps that is acceptable before the first "stable release" of a microservice.
To a small team that doesn't have the inertial issues to generate the benefits of microservices, it seems like they are nice in theory but have too much overhead to supplant monolithic approaches.
Same as any other project: Develop from the outside in.
In practice, trying to develop in the "optimal order' leads to speculative development that will be wasted.
That said, the "it's just like the web" model doesn't sound fantastic to me. It sounds like your app now depends on contracts which are only enforced by good practices, not by something strongly typed you can check at compile time, unless you use something like protocol buffers to generate the boilerplate.
In microservices, everything is dynamically typed.
There is no single binary produced by a single compiler performing whole-program checks of consistency. Even tools like protobufs don't help when code bases drift, or someone introduces a foreign tool, or someone upgrades versions and introduces a subtle mismatch, or some doesn't know you call their service and shuts it down ...
Turns out that driving from tests, and starting those tests from the outermost consumer, is a fairly well-proved way of coping with such conditions.
Static typing is not a panacea, but large codebase plus dynamic typing everywhere sounds like a recipe for disaster. No matter the amount of testing.
> Turns out that driving from tests, and starting those tests from the outermost consumer, is a fairly well-proved way of coping with such conditions.
You need tests no matter what. However, static typing means a much greater confidence in your codebase.
At runtime you are inspecting incoming messages and then routing them to code. It doesn't matter what language the code is written in, it will need to route and validate the messages at runtime.
The type system cannot provide compile-time assurances of behaviour, because it cannot create a single consistent binary which enforces the guarantees.
Your only remaining tool is to drive code from tests and only from tests.
You have serialization/deserialization issues. You can still type your messages.
> At runtime you are inspecting incoming messages and then routing them to code. It doesn't matter what language the code is written in, it will need to route and validate the messages at runtime.
Of course.
> The type system cannot provide compile-time assurances of behaviour, because it cannot create a single consistent binary which enforces the guarantees.
If you make the assumption that you deploy up-to-date binaries, then knowing at compile time that your producer and consumer use the same data structure for the messages they exchange would give me much better confidence than "it looks like the API conforms to what's written on the wiki".
You can hope that they respect the type. For a robust distributed system, you will have to check everything at runtime.
> If you make the assumption that you deploy up-to-date binaries, then knowing at compile time that your producer and consumer use the same data structure for the messages they exchange would give me much better confidence than "it looks like the API conforms to what's written on the wiki".
My reading is that we agree that running code is the only source of truth, we disagree on what guarantees distribution deprives us of.
You're right that tests don't make Byzantine failures go away. But neither do static types. My point that distribution turns all systems into analogies for dynamic language programming remains, and so the emphasis on tool support changes along with it.
In the normal case of development, you tend to have a broken-out system. For a game for example:
+ Game code
--+ User authentication classes/functionality (which accesses DB)
--+ Messaging classes/functionality (which accesses DB)
--+ User metrics classes/functionality (which accesses DB)
In the new design you'd have this: + Game code
--+ User authentication classes/functionality (which accesses REST service)
--+ Messaging classes/functionality (which accesses REST service)
--+ User metrics classes/functionality (which accesses REST service)
In other words, in a clean design, your Game code is accessing a library which provides user Authentication functionality, one which provides Messaging functionality, and one which provides Metrics functionality.In this new design, you have exactly the same thing - a library which abstracts the details of communicating with the service, encoding data, etc. A person making changes to those libraries, which other services use, is responsible for either not making backwards-incompatible changes, or, when that isn't possible, working with other teams to ensure a clean upgrade path (or doing it themselves, if your lines are sufficiently blurred).
The new design trades a "modular but monolithic" design for complexity and brittleness, IMHO. The ability to spin up new instances of a given service on demand is interesting, but it sure sounds like reinventing Erlang without Erlang's tooling.
The only difference is that instead of starting to write some Identity class and use it in your Game service, you write some Identity class and expose it via a REST API, and then provide an interface library that interfaces with that REST API. Call it IdentityInterface or libidentity or something. Pydentity, whatever. It makes an HTTP request, gets a serialized object, unserializes it, and returns it.
For simplicity, put all your public models in that library, and it gets shared by both the Identity service and the Game service. Those models represent an object and what you can do with it. In the Identity service is where all of that actually happens.
This is also how you solve the 'multiple microservices at the same time' problem; your interface library provides the public interface, and the backend REST API is the 'private' API used by the public interface. You make changes to the backend API and the public library and no one notices, or you make incompatible changes to the public service and fix everything before you deploy; ideally, you add new APIs, migrate services over, then deprecate the old ones.
In the end, each service sees the world as fundamentally the same; there's a library with classes and functionality, and you use that to do things. If you design it right, it's never obvious from your code that you're accessing a different service elsewhere in your infrastructure.
If regular object oriented programming languages had method calls that randomly failed, were delayed, sent multiple copies of a response, changed how they behaved without warning, sent half-formed responses ... then yes it would be the same.
Distributed systems are hard, because you cannot change things in two places simultaneously. All synchronisation is limited by the bits you can push down a channel up to, but not exceeding, the speed of light. In a single computer system this problem can be hidden from the programmer. In a distributed system, it cannot.
Probably the most devastating critique of the position that "it's just OO modeling!" came in A Note on Distributed Computing, published in 1994 by Waldo, Wyant, Wollrath and Kendall[0]:
"We look at a number of distributed systems that have attempted to paper over the distinction between local and remote objects, and show that such systems fail to support basic requirements of robustness and reliability. These failures have been masked in the past by the small size of the distributed systems that have been built. In the enterprise-wide distributed systems foreseen in the near future, however, such a masking will be impossible."
[0] http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.41.7...