Good developer experience increases self-reported productivity. I do suspect that it correlates with actual productivity and business outcomes, especially around reducing the productivity loss that comes from error remediation.
Good developer experience increases self-reported productivity. I do suspect that it correlates with actual productivity and business outcomes, especially around reducing the productivity loss that comes from error remediation.
I’ve worked at a lot of big tech companies that do surveys about internal tooling and every year it’s rated as a weak spot, across years and companies this seemed like a consistent trend.
And yet everyone had teams dedicated to improving various aspects of devex so it’s unclear if these teams are just improving the wrong things or if productivity really is improving and it’s something else (eg the amount of code debt grows faster than devex improvements or people are asked to go faster than the devex improvements can keep up or the devex is being improved but the size of the survey means not enough people feel it because you optimize smaller subsets of engineering orgs).
That’s another thing to be mindful about large scale and small scale surveys - the latter might be sampling specific teams adopting the tool whereas the former might find there’s no way to make everyone happy and it all turns into a wash.
You can't measure developer productivity objectively, assuming you're referring to metrics like lines of code, number of pull requests, or velocity points which are infamous. There's broad agreement on this both within the research community as well as practitioners at leading tech companies.
Here are some examples: https://newsletter.pragmaticengineer.com/p/measuring-develop...
"... and what distinguishes 'good devex' from 'bad devex'"
Yes to this. This is an ongoing effort - we have two previous journal papers that touch on this which may be of interest to you:
- https://getdx.com/research/devex-what-actually-drives-produc...
- https://getdx.com/research/conceptual-framework-for-develope...
No those are metrics you’re suggesting. There are better ones as someone else mentioned (time to get code from PR to production, some way of measuring the quality of work getting to production, etc). Yes, the obvious metrics are poor and better metrics are difficult to measure and quantify. And obviously no single metric is going to capture something as multidimensional as code development.
Also, the link you reference doesn’t support your argument.
> Across the board, all companies shared that they use both qualitative and quantitative measures
Throughout it discusses that people do use quantitative metrics to help guide their analysis and none of them try to do the obviously naive ones as mentioned.
This isn’t intended as a critique, but as an engineering profession anything that isn’t quantifiable means it’s open to interpretation and argument weakening forward progress to be restricted to what becomes adopted as industry standard which is in many ways more of a popularity contest of fads rather than concrete technical improvement.
* The time it takes for completed code to be deployed to production * count of manual interventions it takes to get the code deployed * count of other people that have to get involved for a given deploy * how long it takes a new employee to set up their dev environment, count of manual steps involved
Stuff like that
In most places I've worked at, the a survey asking for specific pain points gets great results, because the worst time sinks stick out like a sore thumb, especially if you have workers that have worked in high quality organizations.
There are three problems. The first is that people who can accurately point out the problem are outweighed by a bunch of people who are just unhappy and generate a response for the sake of participation / being prompted. This means that you can address the pain point you think you’ve identified only for sentiment to remain unchanged so you try to tackle the next point and the cycle repeats.
The second is that things like “slow compile times” may have hundreds of different reasons so if you improve compile times in aggregate by 20% you’ve not solved anyone’s specific day to day compile time pain point where they normally spend many 10s compiling which sees a reduction to 8s but the expensive compile which runs infrequently (both because less needed or because it’s so slow) takes 1-4 hours. It’ll maybe see a reduction into the 48min-3.2 hour range which is substantial but not enough to be felt or it may be unaffected because the improvements aren’t measuring that slow build and it’s not a focus (eg maybe it has a bad dependency chain that pulls in way too much). The causes of why that’s slow can be hard to tease out correctly and engineers are incentivized to make the “biggest bang for the buck” changes that sound impressive and quantifiable (20% reduction in compile times across the company vs I made this team happier and it’ll maybe show up in an end of year survey if I’m lucky)
The third is that the rate at which certain kinds of things get worse (long CI and compile times) keeps up or usually outperforms the pace at which things get better (eg 1000 developers adding to compile and test times cannot be beat by a team of 10 engineers spending their time on speeding things up).
Having done internal developer & analyst tooling work (and used DX), this type of survey is great for internal prioritization when you have dedicated capacity for improvements.
I'd be curious to see more about organizational outcomes, as this is piece of DevOps/DevEx data that I feel is weakest and requires the most faith. DORA did some research here, but it's still not always enough to convince leadership.
I expect some HN readers will balk at the idea that we can know things but can't prove them scientifically but only because they haven't really thought about it. E.g. does giving variables sensible names reduce bug counts? Obviously. But it's very difficult to prove even something as obvious as that.
So, we keep trying to come up with objective measures of a term that is impossible to objectively measure. Isn't that weird?
Why not "developer happiness", or, if that's too hippy-dippy for the suits, "developer satisfaction"? Surely we can reason that, all other things being equal, happier developers are more productive.
Then, surveys like this one make perfect sense. "Hey, how satisfied are you with the tools / processes / whathaveyou at your employer?" That's a valid question, and collecting a pile of answers does start to yield some useful insights into what works and what doesn't.
But no, we're trying to use "developer productivity" instead, as though we're all simple machines that provide value by converting coffee into code, and the winning business is the one that is able to extract the most SLOC out of the cheapest and most miniscule quantity of coffee.
It just means how productive developers are. How much they get done. I don't see what's peculiar about that.
Is it peculiar that "productivity" is a term?
> provide value by converting coffee into code, and the winning business is the one that is able to extract the most SLOC out of the cheapest and most miniscule quantity of coffee
That is pretty much the case, though you don't need to word it so awfully.
We provide value by spending our time to producing useful software. The winning business model (all other factors being equal; which they never are) is the one that is able to help developers waste less time when producing said software.
See now it doesn't sound evil.
Explain to me what makes a developer "productive". What, specifically, are you referring to when you define it as, "How much they get done"?
Productivity is how quickly they do those things.
If I hire two developers to make me a website, and they both make websites of similar quality but one gets it done in half the time then he is twice as productive. It's not complicated.
Measuring productivity is of course extremely difficult, hence my original comment about the difficulty of proving productivity improvements. But defining it is not especially difficult.
How is it that productivity as a concept is "not complicated" but also extremely difficult to measure? If productivity has a simple definition, then it should be simple to measure, right? Likewise, doesn't the difficulty in measuring it suggest that maybe it's equally difficult to define?
I am a team lead for an organization; we have 100 tickets in our issue tracker. I have three developers. In the same one-week period, Developer Alfred closed a single issue that had been open for over a year and involved a number of core systems that were interacting in a subtle way and causing intermittent outages; Developer Beth closed 5 issues that were all feature requests from paying clients; and Developer Charlie closed 50 issues that were all small bug fixes.
Which developer was the most productive?
I'm not asking for a measure here. No absolute value or unit of measure required. This is purely comparative. If "productivity" has any meaning at all -- and especially if it's "not complicated" -- then this should have a straightforward answer. How should I, as the team lead, evaluate the relative productivity of my three developers?
I should have been more specific really: it's difficult to measure for (most) programming. Productivity is work/time, and time is obviously easy to measure, but for programming the amount of work involved in a task is really really difficult to know because tasks vary so much.
If you're just building widgets in a factory and they're all exactly the same then productivity is trivial to measure because the work to produce each thing is the same.
But that's not the only reason it's near impossible to set up productivity experiments. Ideally you want two identical people to do the exact same task with only one variable differing (e.g. do they have an IDE or Notepad). But you can't do that. People vary hugely in skill, and you can't have the same person do the same task twice in a row because they'll obviously be quicker the second time.
> Which developer was the most productive?
Very difficult to know because - as I already said - measuring productivity is extremely difficult. Does that mean that productivity is somehow vague or difficult to define? No, absolutely not.
Also... I said it's extremely difficult. That doesn't mean you can't do it at all. Surely you have worked with people who are bad at their job and don't get anything done? Or maybe you're lucky enough to have worked with one or two 10x developers? (It's not a myth; you just might not have met any.)
Yes, I get that the idea of programmer productivity seems simple. It's just how much "stuff" they do! But, once you start trying to work out the details of what "stuff" is and how to evaluate programmer productivity, it all suddenly turns into a big hairy pear-shaped mess. To me, this makes the concept of programmer productivity pretty darn useless.
And yes, I've worked with or observed a pretty wide variety of developers. As is the case with any other industry, there are some geniuses and other people that are just fantastic -- but also, they tend to be people who have found the right role for their personality and interests. I've also seen a lot of fair-to-middling people who had unsung talents but were stuck keeping the lights on somewhere. I don't think any of this is helpful for understanding the meaning of productivity.
Managers have been trying to figure out how to evaluate developer productivity for at least 25 years -- that would've been about the time I read my first trade rag. All they've managed to do is generate a lot of human misery. You'd think 25 years would be long enough for people to come around to the idea that maybe this developer productivity thing they're chasing is an illusion. But ... naaah.
This is the crux of the problem, I think, programmer output can be analyzed in so many different axes that have different weights in different industries that measuring productivity objectively is incredibly complex.
So do ineffective DX, most times. There's "proper" way to assess productivity gains that don't rely on subjective self-reporting (eg the DORA metrics). The "improving feedback loops" item on this report makes sense on that light - the other two are questionable as a direct proxy
I don't find this hard to believe, but do you know of any research showing this?
I've been to multiple organizations that invested heavily in DX, and as an external observer you could see they became actually slower (eg one went from test-in-prod and deploy-to-prod to a convoluted CI pipeline) - it made engineers feel like they were doing a lot less work (you just push code and forget), but in reality, an entire department was hired to deal with issues in prod, and the teams didn't really deliver anything faster at all. But 100% of the devs swore it was better/faster/more efficient.
On another, the added free time converted directly into more complexity, so tons of overengineered microservices (sometimes to the tune of 2-3x the number of individuals on a team) ensued - it got so easy to do X, everyone overdid it.
I generally agree with your point, with the caveat that metrics like DORA are the first to be gamed when applied to an organization.
It's the usual cycle of applying the metrics first in an informal way, then using it for team/process evaluations as it seems to be somewhat reliable, then teams start to reajust how they make the tickets and the deploy cycle in ways that dramatically increase the metrics while reality hasn't changed much.
Analyzing productivity with completely objective metrics will I think always bear these flaws (I'd assume even measuring things at the highest level as "wages vs revenue" would be gamed on how wages and revenue are calculated ?)