Microservices – Combinatorial Explosion of Versions
worklifenotes.com
worklifenotes.com
From the Pact docs: 'Contract testing is a technique for testing an integration point by checking each application in isolation to ensure the messages it sends or receives conform to a shared understanding that is documented in a "contract".' By focussing just on the messages we get tests which are fast, give us quick feedback, and scale linearly instead of combinatorially.
Some good resources are:
https://pact.io (for information about contract testing and the Pact tool itself)
https://pactflow.io/how-pact-works/ (explains how Pact works)
https://docs.pact.io/faq/convinceme (answers the question of why you would want to do contract testing)
https://slack.pact.io (a friendly 1000+ member community which is very experienced dealing with these kinds of issues)
https://docs.pact.io/pact_broker/can_i_deploy (addresses how we handle and channel this combinatorial explosion for good instead of evil!)
The organizing principal was that if my code passed the acceptance tests once, the very same code should still pass with your service/library/data added to the mix. If they don’t, it’s likely your code and not mine.
This might not seem like much but it cuts hours of finger pointing out of a three party system. I expect this effect would be multiplied for a five or ten party scenario.
And in particular, latency has a huge negative impact on behavioral modification. People will keep doing things that they get yelled at over three weeks later. They will stop doing things they get called on hours later.
Which I would hope offsets the additional test matrix complexity.
Anyone who designs a microservice environment like this should be fired. The interdependencies between services in the picture make the entire system fragile and everything will fail if a single service goes down.
A real production microservice environment should be designed so that interdependencies are limited and system wide failures won't occur if a service or two go down. Once you limit the interdependencies then the combinatorial explosion doesn't exist anymore. You might have some services that have a wide range of version used by clients, but you don't have the interdependency that complicates things.
Then, versioning isn't such a big deal.
The dependency pattern around microservices is either a contract on a queue/stream/topic or a version of an API from another depened service.
Don't care how many times you tell me that services should be interdepended and bla bla bla... In the end a service has dependencies and those dependencies must then have a versioning strategy so humans are not in doubt.
Versioning a microservice is no different than versioning a linked library or a Rest API.
Documentation is key. Tell people what versions break.
You can distinguish the versions by using different endpoints, version fields, version headers, etc.. But that's just implementation details.
Otherwise, we're back at "works on my machine" mentality.
And yes - trying to make microservices independent from each other is also a form of pruning and sometimes works well to a point, but also requires good tooling to be done right.
That's a strong statement which you should at least try to provide counter-example to support - just saying "you design it wrong" is not enough. I've seen this kind of problem irl left-right-and-center - so I take it there is an issue, and it's not happening because "everybody doing it wrong".
> If one service goes down, they all go down
I believe you don't understand the diagram. Lines are not dependencies, but a way to connect components of different versions into single product. (I.e., you pick either v1 or v2 of each component, and it becomes single product along the lines - it doesn't necessarily mean there is hard dependency).
To my point, I treat whole architecture as a product. I don't necessarily speak about dependencies. Instead, I'm taking general view of this as a math problem - first of all, I establish search space - that is number of available versions to the power of number of microservices.
Then I very much support and want to discuss various ways to reduce this search space via pruning - and what you're talking about is just one of the options how to prune it (by reducing dependencies). But as others pointed (rephrasing in my words), you can apply a greedy algorithm to NP-problem, but first of all understand what the real problem you're dealing with is and second of all realize that your algorithm is greedy - meaning it may have flaws in the edge cases and it's better to be prepared to those.
In your specific case I claim that you can't be completely sure that you actually don't have any interdependency between components - I believe it would be impossible to prove for any system. And again I saw really hard bugs irl coming from those assumptions. Recent example (this is not about microservices - but pls try to solve this case): https://stackoverflow.com/questions/60486853/aws-ecr-uploadi... - Tools in question must be completely independent, but somehow they are not. So assuming something is completely independent is frequently dangerous.
It doesn't.
No matter how few or many dependencies you have, when they are updated, the dependent system must either update or not update.
There is nothing new in regards to versioning of dependencies. It has always been hard to do, and with microservices the versioning story has not all of a sudden improved.
What people that are new to versioning, seen from a somewhat autonomous systems perspective, this is something you must have a strategy for.
This sounds crazy, but it's essential if you can't guarantee that you can atomically deploy a new service version AND any other services that call it.
Creating backwards compatible releases means sticking to some rules, things like:
- you can add fields but you can never delete them
- you can add new methods but you can't remove old ones
Having really detailed logging helps a lot. If you want to deprecate a method you can remove it from calling clients first and then use the logs to confirm it isn't being called any more before removing it from the service.
It's a lot if work, and teaching a large engineering team how to be productive in this kind of environment is decidedly non trivial.
Adding fields or methods is easy. You can do it at any time.
Removing them is difficult. You must first go through a long deprecation cycle and ensure ALL clients are updated.
You need good logging AND good monitoring and analysis of the logs in both client and server to be able to detect access to removed features.
Alternatively you could provide different API versions at the same time and keep the old ones running, as long as they are called.
And, in general, don't focus on generic methods.
The more specific a service method, the better. The less it does, the better. Don't have an "order pizza" api. Have a "start order", "add to order", and "complete order" api. If it helps, have multiple add methods. The parameters to add a pizza can and should be different from a salad. Even the pizza may be better served with partial service calls to start and complete.
Of course this does not "solve" the problem of incompatibilities themselves - for that the simple solution that I've seen used is just bog standard versioning of APIs (/v1/... /v2/... /v2019-03-04/... etc) and some procedures for managing the supported versions, e.g. not doing breaking changes in an existing version, only supporting a set number of versions (i.e. previous, current, and next), and having proper sunset-periods for older versions before they are turned off so clients have time to update. Slows velocity to do all this management but that is life anyway in a complex system
But in reality, you're still facing same issue - let's say your components X and Y are updated in monorepo. You want X in production, but not Y - so now you have to invent something on top of monorepo to capture this. Another words back to square 1.
Regarding startups, in my experience they start feeling this pain when they already have 5-6 microservices. But YMMV - depends on a lot of factors.
1. Always be backwards compatible in protocols and data formats.
If you store data on disk or in a table you V1 and V2 tables need to be generally compatible.
V2 should always be able to read V1 tables. This helps with rollouts.
In an ideal case V1 should gracefully handle V2 tables. This helps with rollbacks.
2. Service to service communications work between V1 and V2.
If you don't want downtime you have to gracefully handle the situation in which V2 is rolling out. In a zero-downtime deployment config like blue/green V2 instances can get requests from V1 instances and vice versa.
Monorepos don't help with these situations. These are just engineering practices that have to be upheld by the team. Monorepos will make the atomic source code updates applied across the codebase easy.
A startup should be using a monolith until it becomes a victim of its own success and then should migrate thoughtfully to microservices only when it has to. It also requires a very heavy investment in devops and tools since microservices will fail in production in every which way possible. It also requires a huge investment in metrics and alerts and logging otherwise you will have no idea what is wrong with your now too-complicated production system.
While I generally agree with you in the vast majority of practical cases, it can make sense in some circumstances, namely, if there is a very clear logical abstraction between one service and another.
A microservice is just another layer of abstraction, similar to functions, classes, files, libraries, and programs, but even higher up the food-chain, and it comes with its own characteristics and peculiarities. If you have high confidence that one component is (1) logically extremely different from all of the other components, (2) has no side effects, (3) is not I/O bound as it's going over a network, and critically, (4) is on a separate development timeline, then you have a solid candidate for a microservice.
In my experience this is not always the case and the problems creep in in other places that are not easily statically checked or tested, but which come up at runtime, e.g. specific runtime content/data, or serialisation of some message/object that happens at runtime.
1. Stop all old servers before starting the new ones -> downtime, might be acceptable 2. Have your clients handle the fact that they might hit vN or vN+1 of your services.
Note: None of these are cast in stone. These have worked well for us
1. Decide in your team/org what constitutes a new version and what it means. Does every change means version change? What if an API adds an optional parameter; do I need to update the version?
2. Not everything has to be a service. Think if the same functionality can be consumed in the form of a library/sdk instead of a full service.
3. The dependency amongst services should resemble acyclic graph and not cycles. This limits the version change impact.
4. Think about abstraction. Can we logically group set of services A, B & C and provide unified API via pass-through service D? Only D needs to handle API breakage most of the time.
If each service can talk to each other, the number of possible communication paths grows exponentially -- this is manageable up to a point, but even the largest teams will have to start investing into orchestration sooner or later.
There are many orchestration patterns and mechanisms, and this is not the place to list them all -- which one is the right choice will depend on the complexity and needs of the system.
And yes we can agree that there are various (usually "greedy") strategies to simplify this - but this problem is a fact of life and it makes thinking process easier to accept it as such. Same as accepting that it is impossible to solve consensus in async system, but there are algorithms that perform fairly well.
+ I summed this up in more details in the comments down below.
There are good reasons to use microservices, like when you have 100+ engineers working on your planet-scale cloud or social network. At that scale, you need to have internal API documentation to ensure people that never met can still work together productively.
But if you are a normal company, you will probably never have enough engineers to make microservices worthwhile in the first place. So just put everything into one big git repo and call it a day.
Also, the ability to do rolling restarts and no downtime upgrades by running multiple versions in parallel adds a lot of work and complexity. But for most small to medium companies, a planned 5 minute downtime in the middle of the night is completely no problem, so all that zero-downtime-upgrade work is just wasted effort.
I mean, even my bank has a fixed offline maintenance window every night from 3:00 to 3:30 am. As does Amazon RDS.
So don't solve problems that you don't have :) Most companies do fine without microservices.
But I agree wholeheartedly. You don't need 100+ services to get the benefits of converting your tool chain to handle deploying n services even if n=3.
Like every other fad in technology the nuances and practicalities are lost along the way.
Now that we have all started bashing microservices can we all take a minute to reflect on what a steaming pile of shit monoliths can turn into without a huge amount of respect for the artificial boundaries and interfaces that you create.
Neither is really there, and so I agree with you today, but it could change with the right tooling.
The first half of that is true, but the conclusion overlooks a couple of much bigger issues:
1. There are few indicators of effective software engineering so strong as deployment frequency and lead time for changes. Being able to go from issue to fix in production in minutes rather than days opens up completely new paths for feedback and allows you to operate your service much more efficiently. (See the yearly state of devops reports for further research about this.)
2. Not all downtime is planned downtime. Being able to handle a true crash gracefully is very important to achieve reasonable response times in the high nines.
I'm not a proponent of service-oriented architectures -- my opinion is that enforcing API boundaries is useful, but can be done at the source level rather than application level. I do however think there is great value to being able to deploy/restart your application with virtually no downtime.
This is only true under certain assumptions. E.g. if your business is not a public facing website, you often don't have the immediate feedback in case of errors, to make the above observation applicable. I have a customer with an application that gather date for a whole year, and then does alot of stuff with it. Then, there are no short feedback cycles in production, no matter how often a day you deploy. Then, it is important to assure quality before a change reaches production. Otherwise it might replace your data with background noise and you only discover it at the end of the next year.
This is a symptom of the general tendency in developer circles to assume that personal findings apply to everyone at every time in every situation. Which is obviously ridicoulous. But somehow completely acceptable if you generalize from something one Google employer has posted about.
Even if the latter category is more numerous, I think we spend much less time working on it. And software that's used semi-frequently can benefit from a shorter feedback cycle, even if it doesn't have one right now. Sometimes there's an insurmountable technical barrier to that (e.g. code on a solar orbiter) but it's even in those cases worth investigating.
3:00-3:30 in your timezone ;)
You don't need (micro)services to do upgrades without downtime. You just need a couple of load-balanced servers instead of one, and some planning to be sure that the two versions can coexists for some time.
This can be a bit tricky until you get used to it, but depending on the company and the customers it can be even easier than organizing a maintenance window.
Its enough to have 2 that work in different domains with different constraints and different tools. Company scale is completely irrelevant here.
> But if you are a normal company, you will probably never have enough engineers to make microservices worthwhile in the first place
Many "normal" companies use microservices and common reason is that they simply cannot function otherwise.
> But for most small to medium companies, a planned 5 minute downtime in the middle of the night is completely no problem,
The fact that some site doesn't function for 5 minutes may not be a problem, the fact that any change requires many people agreeing on deployment time often is.
Wouldn't that be more or less be just two monoliths?
If they are completely isolated, than yes, but its really rare scenario.
A common example would be a website that gathers/presents some data from users, and a service that provides analytics of said data. Both domains may require completely different technologies.
Surely this doesn't actually happen.
Imagine
a - you have 10 microservices, 1 update per each. Each supposed to be backward-compatible. You start rolling out. One by one. 5 go fine, 6th breaks. You end up in a weird state where 5 out of 10 are updated.
You hope it's fine due to backward compatibility but you never really tested this config exactly. (Which is why I might prefer either converging to deploying all 10 or rolling back fully - if that was a known good state - and having canary cluster rather than canary microservice in many cases).
b - same as above but now you try to catch this behaviour on test / staging. You still have same hard problem at hands.
Key here is you clearly can't try every possible variation of what may break, so need to make conscious decisions about what to do in the case of failure.
So in many case, you will have only a single version per service.
Moreover, imo, when we are talking about internal apis, there should be nothing preventing the service owner from updating the consumers if he wishes to converge more quickly after breaking compatibility - just like he would when refactoring a monolith. The culture should allow and encourage this kind of collaboration.
Making he interface of services backwards compatible, either by testing and/or by using stricter contracts like grpc (which had a way to deal with compatibility)
Then you're mostly covered. However integration testing indeed makes sense
Microservices should not be mass deployed together but be managed separately. That vastly simplifies the combinatorial explosion of versions you need the worry about: the currently live ones.
The rest is just applying SOLID principles to your microservices and avoid having services with poor cohesiveness or tight coupling (most of the SOLID principles boil down to affecting these two metrics). Bad service design with tight coupling where deploying service A also requires redeploying B,C,and D, is basically a design problem and not a micro services problem. There's a lot of bad design in our industry. Monoliths allow you to get away with that but it's also the reason that breaking them up is a hard problem. Just because you are getting away with it does not mean it is not a problem though.
You can further mitigate integration issues by deploying using modern practices like blue green deployments where you gradually move traffic to a new versions and adapt on things like error rates and other metrics, AB testing, etc. So if it breaks, you don't end up breaking it for everyone and you can fix the problem and try deploying that in a controlled way.
Basically your goal is to keep your customers (aka. dependees) happy and make sure you don't negatively affect them (which you should be actively monitoring). Likewise if you have dependencies (to whom you are a customer) and they stop working when you deploy something new, you probably want to detect and fix that before you break all your customers. If that is a regular thing, consider having integration tests, contract tests, etc.
Staging environments are a controversial topic in this context. My view is that they don't make much sense in a properly run micro service deployment since you are not testing in a real environment with real users, real data, and lots of things happening concurrently and features interacting. And of course your customers doing real things they care about. The only realistic environment that has that in most complex micro service deployments is called production. Once you add serverless and edge computing to the mix, these things become even more true. The bigger the organization, the less feasible it becomes to have a staging environment.
Update. I forgot to add this but doing continuous deployment means small deltas that are low risk. Any bigger change can be behind a feature flag. There are a few more strategies.
Also, I'm not actually a microservices proponent for small teams. It just creates deployment and operational overhead (read go to market bottlenecks).
Staging envs are generally a crutch. They can help but they can also hurt.