For teams you can measure meaningful outcomes and improve team metrics.
You shouldn’t really compare teams but it also is possible if you know what teams are doing.
If you are some disconnected manager that thinks he can make decisions or improvements reducing things to single numbers - yeah that’s not possible.
How? Which metrics?
I think it's harder to measure things like developer productivity. The closest thing we have is making an estimate and seeing how far off you are, but that doesn't account for hedging estimates or requirements suddenly changing. Changing requirements doesn't matter for DORA as it's just another sample to test for deployment.
Unfortunately there's a lot of lag
A great generalisation and understatement! Often looking like you are becoming more efficient is more important than actually being more efficient, e.g you need to impress investors. So you cut back on maintenance and other cost centres and the new management can blame you in 6 years time for it when you are far enough away from it to not hurt you.
Then I would have to go on outlining 6-12 months of trying stuff out.
Because if I just give "an example" I will get dozens of "smart ass" replies how this specific one did not work for them and I am stupid. Thanks but don't have time for that or for writing an essay that no one will read anyway and call me stupid or demand even more explanation. :)
This seems to be useful to understand and internalize that there are no simple answers like "use story points!".
There is also loads of people who don't understand that, so I stand by that is useful and important to repeat on every possible occasion.
Measuring it is not the hard part.
The hard part is doing anything about it. If you can't attribute specific outputs to specific inputs, you don't know how to change inputs to maximize outputs. That's what managers need to do, but of course they're often just guessing.
But I do code myself, I write requirements so I do know which ones are trivial and which ones are not. I also see when there are complex migrations.
If you work in a group of people you will also get feedback - doesn't have to be snitching but still you get the feel who is a slacker in the group.
It is hard to quantify the output if you want to be removed from the group "give me a number" manager. If you actually do the work of a manager so you get the feel of the group like who is "Hermione Granger" nagging that others are slacking and disregard their opinion, you see who is the "silent doer" or you see who is "we should do it properly" bullshitter you can make a lot of meaningful adjustments.
Even that would be hard since hunting is complex. If you are the one chasing the pray into the arms of someone else, you surely want it to be considered a team effort.
You need like 'blueberries picked'.
"You can [accurately and meaningfully measure software engineering productivity] - but not on the level of a single developer and you cannot use those measures to manage productivity of a specific dev."
At the level of a company like Google, it's easy: both inputs and outputs are measured in terms of money.
I am not Amazon person - but from my experience 2 pizza teams was what worked and I never implemented it myself just what I observed in wild.
Measuring Google in terms of money is also flawed, there is loads of BS hidden there and lots of people paying big companies more just because they are big companies.
So that's how animal husbandry came about!
The superstar developer’s secret… he would send blank reports to clients (who would only realize it days later, and someone else would end up redoing the report), and he would score many more points without doing anything. I’ve seen this happen a lot in many different companies. As a friend of mine used to say, “it’s very rare, but it happens all the time.”
I have no doubt that AI can help developers, but I don’t trust the metrics of the CEO or people who work on AI, because they are too involved in the subject.
1) They can work to improve the system
2) They can distort the system
3) Or they can distort the data
ah, to be young again...
Now, I was really bad at capitalizing on it, so nothing much came of it, but still, there are some positive things that higher-ups do notice.
None of this works to evaluate individuals or even teams. But it can be effective at evaluating tools.
To use your example, a user with an LLM might say "LLM please fix this" as a first line of action, drastically improving this metric, even if it ruins your overall productivity.
edit: typo
That is, I, personally, am not measured on how much AI generated code I create, and while the number is non-zero, I can't tell you what it is because I don't care and don't have any incentive to care. And I'm someone who is personally fairly bearish on the value of LLM-based codegen/autocomplete.