> that **false** proxies are game-able.
You say this like there are measures that aren't proxies. Tbh I can't think of a single one. Even trivial.All measures are proxies and all measures are gameable. If you are uncertain, host a prize and you'll see how creative people get.
https://en.wikipedia.org/wiki/Net_present_value#Disadvantage...
Second off, you don't think it's possible to hide costs, exaggerate revenue, ignore risks, and/or optimize short term revenue in favor of long term? If you think something isn't hackable you just aren't looking hard enough
But I'm suspicious of a claim that it can't be hacked. I've done a lot of experimental physics and I can tell you that you can hack as simple and obvious of a metric as measuring something with a ruler. This being because it's still a proxy. Your ruler is still an approximation of a meter and is you look at all the rulers, calipers, tape measures, etc you have, you will find that they are not exactly identical, though likely fairly close. But people happily round or are very willing to overlook errors/mistakes when the result makes sense or is nice. That's a pretty trivial system, and it's still hacked.
With more abstract things it's far easier to hack, specifically by not digging deep enough. When your metrics are the aggregation of other metrics (as is the case in your example) you have to look at every single metric and understand how it proxies what you're really after. If we're keeping with economics, GSP might be a great example. It is often used to say how "rich" a country is, but that means very little in of itself. It's generally true that it's easier to increase this number when you have many wealthy companies or individuals, but from this alone you shouldn't be about to statements two countries of equal size where all the wealth is held by a single person or where wealth is equally distributed among all people.
The point is that there's always baked in priors. Baked in assumptions. If you want to find how to hack a metric then hunt down all the assumptions (this will not be in a list you can look up unfortunately) and find where those assumptions break. A very famous math example is with the Banach-Taraki paradox. All required assumptions (including axiom of choice) appear straight forward and obvious. But the things is, as long as you have an axiomatic system (you do), somewhere those assumptions break down. Finding them isn't always easy, but hey, give it scale and Goodhart's will do it's magic
Exactly (btw. very nice way to put it)
> stop wasting time on tracking false proxies
Some times a proxy is much cheaper. (Medical anlogy of limited depth: Instead of doing a surgery to see stuff in person, one might opt to check some ratios in the blood first)
This would not count as a false proxy however. The problem in software is, it is very hard to construct meaningful proxy metrics. Most of the time it ends up being tangential to value.
I agree in principle, just want to add a bit of nuance.
Lets take famous “lines of code” metric.
It would be counterproductive to reward it (as proxy of productivity). But it is a good metric to know.
For the same reason why it’s good to know the weight of ships you produce.
The value in tracking false proxies like lines of code, accrues to the tracker, not the customer, business or anyone else. The tracker is able to extract value from the business, essentially by tricking it (and themselves) into believing that such metrics are valuable. It isn't a good use of time in my opinion, but probably a low stress / chill occupation if that is your objective.
In theory you can return to the metrics later for shorter intervals.
From a programming standpoint, and off the top of my head, I would include TDD, code coverage, and anything that comes out of a root cause analysis.
I tell junior devs who ask to spend a little more time on every task than they think necessary, trying to raise their game. When doing a simple task you should practice all of your best intentions and new intentions to build the muscle memory.
I don't know how to track TDD, but for me, code coverage is an example of the same old false proxies that people used to track in the 000s.
Before creating a metric and policing it, make sure you can rigorously defend its relationship to NPV. If you can't do this, find something else to track.