Measuring developer productivity with the DX Core 4
getdx.com
getdx.com
It is called "not wanting to work in a company that is insulting me with measures like metrics, spying on screens, squeezing every drop of sweat and exploiting its workers for already huge profit.
My last job change was exactly due to company acquisition by USA company that started to exploit everything work related, removing wfh, ruining life-job ballance,... And who leaves first? Those who can.
So you are not really measuring developer productivity but helping company get rid of most productive people, while all others stay.
All I see tools like this doing is getting people to manipulate the metrics, while at the same time producing enormous amounts of worker alienation. I'd even speculate the demoralising effects of workplace surveillance schemes undermine the alleged productivity benefits.
I don't have any data to prove this :)
good
Focus on the stuff that matters, not on tormenting the people that make you money.
there's strong correlation when onboarding an engineer between the time to first/tenth/fiftieith PR and their PR throughput two years later, so the answer is make it really easy to do PRs when onboarding, and they saw the same outcome even if those initial PRs were trivial
through DORA/SPACE/DevEx they also found that asking your engineers if they feel productive is generally as useful and reliable than trying to measure every possible dimension - which is why DX is based largely around the survey with subjective responses. if you actually use those responses to address friction and frustration, from my experience you end up with more, better, easier work getting done
it's possible to work in a team that cares about measuring their effectiveness, and using those measurements to understand how to be more effective. most/all engineers will have an idea of how they could improve their work, but the reality when you actually spend time to measure often shows a lot of different things you might not have been aware of
correlating all of these let's you know if you're accelerating one activity metric but getting worse in an outcome metric - lots if slop PRs but more incidents is obviously the wrong outcome, and can be addressed by improving local verification, safer deployments, faster rollbacks, better observability, etc.
even with AI writing the code (and much more) all of these things matter. actually seeing positive outcomes (and thus RoI from AI use) requires knowing what's signal and what's noise. ive spoken to founders who see anecdotal increases in everything negative as their AI code volume goes up, and none of them are actually measuring anything effectively enough to understand what to do about it
My experience is that these are only useful over a long enough period and across enough people that we can spot genuine outliers. For example, your average across the last 6 months is "15" and the average of everyone else is "25", can you help me see whether you are being given too much off-target work or are there any other issues that are blocking you getting stuff done?
Funny anecdote: I had a large company selling their LLM tools with a slideshow starting with a mention of Goodhart's Law, and then proudly present number of PR's as a metric in the next slide.
That's why there are so many stupid ideas floating around.
I have quit companies that made their performance review systems too onerous/frequent/intense. It’s exhausting to constantly be justifying your existence and having to sell yourself instead of doing the real work.
At one point I was at a 50/50 ratio of time spent documenting/justifying my work vs doing my work, and that was miserable. My level of anxiety reduced my job performance, which created a very unhealthy cycle. Pressure creates diamonds but it also suffocates and crushes most things.
That’s why seeing things like in-context sampling in the article make me suspicious. I don’t think I would be happy at all being interrupted as I’m in flow state to be asked if I could be more productive. I also feel some of this like speed (“time to 10th PR”) or working on new stuff is creating bad incentives.
Fundamentally the most important metric is how hard your boss will fight for you. The second most important metric is whether you can meaningfully discuss areas of improvement honestly and safely with your boss.
If either of those two are out of whack you should question whether to make a change. If I’m missing those I’m not going to be able to do my best work, so it cuts both ways.
My company leadership does not use these as a one size fits all metric. In fact, the code review & PR stats never even show up in performance reviews. Mainly, they are used to improve systems around dev productivity. Like, we see that generally, we’re shipping less PRs and the survey shows people think it’s getting slow to merge changes. That is an opportunity to look at where our systems are holding us back.
I think broad stats are also probably still useful too. For example, there have been times that I’m reviewing more code than everyone on the team combined. I WANT my manager to know that’s happening and figure out how to even the workload. Or, why is my team shipping like 2-3 times as many PRs as a similar team with similar headcount? It’s not immediately a PIP or something, but it’s highly likely something odd is happening. Could leadership improve direction & focus, make priorities more obvious?
When it is decided that a management resource such as a line manager, director, or even C-level exec are causing detriment to organizations they can be put on performance improvement plans —which involves developers simply ignoring anything and everything that person says until they stop speaking nonsense.
This is the way.
I think that’s a good thing, and I say that as an engineering manager.
Having been an IC, I’m acutely aware that developers, when aware of monitoring via performance metrics, will game that system.
Those are horrible incentives.
Sounds like the name of a new CPU/GPU.