McKinsey: Yes, you can measure software developer productivity
mckinsey.com
mckinsey.com
Good stuff for the next bullshit bingo template though, I didn't have "Developer Velocity Index Benchmark" yet.
If I’m understanding them correctly, their pitch is basically “pay us a lot of money to teach you some basic software engineering concepts, and don’t use dumb metrics like LoC produced per day.”
I think paying McKinsey prices to get this is a poor value, but it would probably improve most software engineering management.
That's really key. A company actually interested in technology wouldn't be in the situation of having an engineering team they have no idea how to manage.
So the target here is the kind of company that probably ended there by throwing money around, and would like to throw money at this problem as well. McKinsey will be there whenever there's money to be thrown at a problem.
Their highly sciency formula of number of commits divided by SLOC of the average pull request plus average response time on slack will scientifically segregate your company's developers by productivity.
I do hope we see a follow up on Hacker News - how to game the fuck out of McKinsey's top secret developer productivity formula.
"use our stupid metrics instead"
"Did someone say TDD? Cause I think I heard someone say TDD..."
No more than anything else they charge for.
"Assessing contributions by individuals to a team’s backlog (starting with data from backlog management tools such as Jira, and normalizing data using a proprietary algorithm to account for nuances)"
Using nonsense metrics like "data from [Jira]" apparently.
Though the other columns in the individual row - including examining developer satisfaction, retention, and daily interruptions - might be a slightly more meaningful, since they evaluate the work environment rather than trying to evaluate individual developers.
What real teams and coaches do is say we have a specific way of doing things. Everyone has a role and a person in that role is expected to do X Y and Z. While a different role is expected to do U V and W.
This is what I object to the most to modern tech. It treats technology stacks and developers as interchangeable which is far from the truth. It is this factory mindset that I would argue is creating garbage software.
Ever heard of Agile Sprint Velocity? Yeah the bastardisation of SWE began long before mckinsey floated up to the surface.
The big-picture view is given by Exhibit 1, and you can see it's really much more about analyzing all aspects of productivity (company-level, team-level, etc.) to identify bottlenecks around e.g. too much time spent testing over writing code, or senior developers stuck in endless cross-team alignment meetings which means management needs to identify a single person responsible for making final decisions quickly. So this is a document for engineering managers, not engineers.
But in terms of measuring actual individual developer productivity, if you read through all the verbiage it's really mostly just adding up the "planning poker" agile/sprint points that a developer delivers over the course of months. About which I have two main thoughts:
1) As long as planning poker is executed correctly -- i.e. always with unanimous consensus from engineers and without managers/PM's providing pressure -- this is actually probably going to be pretty decently objective, as a moving average. It's not terrible -- it's infinitely better than things like "lines of code". Although teams can also experience gradual "points drift" over the course of months, and points between teams can't be translated, so it can't be used company-wide or even year-to-year. But it will compare productivity between team members.
2) BUT, there's a huge risk involved because not all e.g. 5 point stories are the same, because of variance. Some are very straightforward where you know it's 5 points, while others are much more unknown, where the best guess is 5 points -- and it might wind up being 2 points but it also might wind up being 50 points. And so I could definitely see certain stories turning into "hot potatoes" nobody wants to touch because it's just too much of a risk -- it'll tank their average and they won't get promoted. And then everything just breaks down. Which means you either need post-sprint processes for "correcting" estimates and assigning a reasonable value for points done, or going deeper into "research" stories -- e.g. a 1-point pre-story to better nail down the actual size.
But story points themselves are actually guesses, and are never a true reflection of the work undertaken post mortem - and rarely if ever get reevaluated with 'correct values' mainly because it's really difficult to do that - it's another guess really.
I think any attempts to measure the accuracy of guesses on imaginary relative units, specific to a team, with deliberately obfuscated and difficult to reconcile units is futility by definition.
Instead companies have started to see the light of measurement using standard units of time and tracking the start to finish time including the states of tickets - Flow metrics.
Which is all fine, however you still have the original problem - unless you strictly compare very similar tasks perhaps done by similar engineers in terms of tenure, skill level and maturity, and maybe other factors, you still can't directly compare tasks across teams.
There still needs to be consideration of the make up of a team to get an understanding of skill level and experience.
Ive personally witnessed teams with graduates in them and people with 20 years experience completing very similar tasks and the difference is astonishing.
Step 1: We are going to estimate tasks in imaginary relative units
Step 2: We are now going to use these imaginary units to measure your velocity
Step 3: We are now going to compare your progress in imaginary units to other teams who use a different imaginary units, and complain your underperforming
Then you imaginary unit inflation and before you know everything is an XXXXXXXXXXXXXXXL t shirt size.
The same people are responsible for the estimate as the delivery. The obvious maximum solution for a team is to continually inflate all estimates to demonstrate improved velocity. At the individual level, it is to always take the most over-inflated story in order to demonstrate maximum individual contribution.
You get perverse outcomes with all metric based productivity measures (for example in retail sales having friends buy large volumes and return to another store). You can prevent this kind of thing to a degree through various sticks, but with “planning poker” it is the antithesis of the desired goals. Measure in a different way.
> leaders and developers alike need to move past the outdated notion that leaders “cannot” understand the intricacies of software engineering
Leaders absolutely can understand software engineering, if time and effort to learn is spent. McKinsey themselves say this. So their argument is 1) if you learn about software engineering, you’ll be more capable of evaluating the efforts of software engineers and 2) don’t use idiotic quantifiers of productivity like LoC per day.
the company has now realised and is now rolling back everything they did
but sadly lost probably 2/3rds of its good developers in the process
No more McKinsey. Minimize management.
In this case, worthless crap.
I’ve since been responsible for tech teams as leader of larger businesses.
My take is it seems basic but sensible.
>generative AI tools such as CopilotX and ChatGPT have the potential to enable developers to complete tasks up to two times faster
I'm surprised McKinsey is willing to endorse ChatGPT, which is one of their major competitors in the information-free drivel market.
I really believe on the team in general can be held accountable as a whole. The team knows who isn’t performing and can ask for them to be removed. If the whole team is not meeting deadlines it should be disbanded.
This is a similar thing. If you have the ability to understand how to interpret the metrics, the metrics can be useful. If you can't interpret the metrics, you will use them poorly and the result will not be a better organization.
Fire 50% of middle management. Use the money saved to hire more developers.
You are welcome.
McKinsey, but for helping companies undo whatever McKinsey recommended to measure software developer productivity. Or just hire them again.