Slow Deployment Causes Meetings
facebook.com
facebook.com
Meetings are caused by policies and the need for coordination. Even companies with continuous deployment still have meetings.
You can eliminate policies, which will reduce meetings, but you'll still need coordination. You can move coordination to your issue tracker or agile board, but then you've just replaced meeting time with the time you spend interacting with the tickets and the agile board.
There really is just no way around it. The bigger you get, the more coordination you need.
Running small teams responsible for small services is one way to reduce that need somewhat, which is part of why microservices are getting more popular, but at the end of the day I'd say coordination time, in whatever form it takes, is linearly (or maybe even exponentially) correlated with org size.
I do think there's an element of truth to what he's saying: organizations develop processes that can handle a certain amount of throughput, and trying to increase throughput means making changes to those processes, which means meetings about reducing the number of meetings. The amount of overhead is largely driven by executive visibility requirements (and past times that executive has been "burned" by a bad deployment) so it can't really be changed.
I actually kind of agree with the mantra: "If you want to deploy more, deploy more often." We become accustomed to testing changes of a certain size, and if we try to cram more into a release, then release testing and certification takes longer. Big releases are the enemy of most Agile processes.
Not if you do it right. As long as the API doesn't change and you adhere to every service having its own data store (which is key and many people forget), you can make as many changes to your service as you want without any coordination.
I'm really not sure that's true. If I change my classifier endpoint, even with the same API it's important people know that they might get different results for the same input. It might be really quite important for peoples analysis.
Their system might not fall over, but that doesn't mean I don't need to coordinate with them.
Perhaps I'm missing something in what you're describing though.
Although that does involve knowing your responses are identical.
If my results are different at all I would need to notify people downstream.
Obviously, I can't start suddenly returning XML instead of JSON, because "I feel like it and it's cool"; it breaks everyone horribly, so that needs to be version 2.
But if you insist on upping the version effectively any time you do, you're essentially saying that any side-effect (not just output) visible can't be fixed, except by duplicating the API and creating a new version, fixing it there, and moving the callers over (if you can — this, I've found is the hardest part[1]).
Often, these are things that I feel like the spec is supposed to hash out: what is it supposed to do. If it strays from the spec, that doesn't necessarily require a new version number, especially if it doesn't break things.
However, it has been my experience in that even the most benign changes will break things; you've got to be able to evaluate and figure out who is not following spec, and sometimes, how bad is the effect of the change if the clients didn't like it. (Is it so bad that, even though it's to spec, I should roll back? Or should the client fix the buggy behavior they've been depending on, and we can carry on in the meantime?)
[1]: Mobile clients are out there. Unlike a web page, there's no way to edit/update them — you can push a new version, but will the users upgrade? At some point, you have to cut them, I suppose, but after how long? 6 months? <2% share among the user base? the active user base? What's an active user? (and now higher ups are involved, and you're in meetings…)
[also]: https://xkcd.com/1172/
Presumably your API says what types of data the caller should expect. As long as you don't change that, they should be able to deal with the response changing.
But really it's more about the fact that if you have a single monolith, if you want to make a change, you have to coordinate with everyone at the company, whereas if you have microservices, you only have to coordinate with those who are affected by the change.
The point I wanted to get across was that although a client won't break, I'd still want to talk to those pulling data from the system. I was disagreeing with the statement "you can make as many changes to your service as you want without any coordination."
> But really it's more about the fact that if you have a single monolith, if you want to make a change, you have to coordinate with everyone at the company, whereas if you have microservices, you only have to coordinate with those who are affected by the change.
Yes, it does lower the hurdles to getting something out.
And very rarely can you implement a feature without touching multiple microservices. Hell, that's why the concept of epics even exists: you have many smaller user stories inside the epic that need to be coordinated to deliver a single feature. Your UX needs the back-end service calls to exist before they can release their UX changes.
Mentally, I imagine an airport and there's a number of planes that need to be landed in a week. What requires more coordination: an airport that is open 24/7 and planes land when they arrive (continuous deployment) or an airport that is only open 1 hour a week?
Congestion drives the need for coordination. Or, put another way, when there is low or no congestion, sequencing through time is a sufficient form of coordination.
But that happens asynchronously, which is already an improvement.
> There really is just no way around it. The bigger you get, the more coordination you need.
True.
> (...) at the end of the day I'd say coordination time, in whatever form it takes, is linearly (or maybe even exponentially) correlated with org size.
It's usually quadratic. If you manage to make it linear then you've probably found an organizational goldmine...
Strict tree chain of command. O(n log n) coordination cost. Much better than quadratic. Approaches linear with the flatness of the org.
(Semi dupe of my comment above, but I couldn't resist)
This is how very large orgs like the military scale. Common sense shows it can't be exponential, since the military is not totally paralyzed.
"Increasing overhead initiates a positive feedback loop: less getting done -> more pressure -> more mistakes -> even fewer changes per deployment -> more overhead -> less getting done."
I've seen places where waterfall was totally appropriate and worked well. There's a lot of variety out there.
The mobile app seems a bit slower, with broken features staying broken for a while. (matching the stated 2 week cycle).
Most production environments are much more constrained and/or operate under a lot of outside constraints. In those situations you have to be a lot more careful.