Like can we determine the productivity of doctors, lawyers, journalists, or pastry chefs?
What job out there is so simple that we can meaningfully measure all the positive and negative effects of the worker as well as account for different conditions between workers.
I could probably get behind the idea that you could measure productivity for professional poker players (given a long enough evaluation period). Hard to think of much else.
The British government (probably not any worse than anyone else, just what I am most familiar with) does measure the productivity of the NHS: https://www.england.nhs.uk/long-read/nhs-productivity/ (including doctors, obviously).
They also try to measure the performance of teachers and schools and introduced performance league tables and special exams (SATS - exams sat at various ages school children in the state system, nothing like the American exams with the same name) to do this more pervasively. They made it better by creating multi-academy trusts which adds a layer of management running multi-schools so even more people want even more metrics.
The same for police, and pretty much everything else.
And to be fair, some crud work is repetitive enough so it should be possible to get a fair measure of at least the difference in speed between developers.
But that building simple crud services with rest interfaces takes as much time as it does is a failure of the tools we use.
Yes, yes we can.
Programmers really need to stop this cope about us being such special snowflakes that we can't be assessed and that our maangers just need to take that we're worth keeping around on good faith.
Of course we can. But can we do it in a meaningful way, such that the metric itself doesn't become a subject to optimization?
"When a measure becomes a target, it ceases to be a good measure"
By making the metrics part of a sustaintable company-wide goal. If there's a company-wide goal to increase X kind of revenue by Y% making actionable targets on how a team can contribute (not lazy shit like "our changes should contribute Z% of that Y%"), and within that create for a person another smaller metric based on that.
Also, medical facilities… you certainly could define it as profit, but that bothers me and many other people.
You could define it as patients seen, or “cured” but that incentivizes very quick but probably poor care.
You could define it as intensity of treatment or amount of care given, but you’d probably end up in a situation where 1 incredibly sick person has every doctor treating them.
You could define it as…
In real world, most things don't work out that way. What metrics do you use to measure surgeons' success? If you use fatality rate, then as a result surgeons will refuse to do more risky surgeries which will put their ratings at risk, which makes the healthcare worse, instead of better.
Could you make an effort to explain how, or at the very least link to some reasoning? Otherwise your comment is basically the equivalent of “nuh-uh”, which doesn’t meaningfully contribute to the discussion.
> Programmers really need to stop this cope about us being such special snowflakes
Which is not at all what is happening in your parent comment. On the contrary, they’re putting developers on even footing with other professions.
You can look at the kind of work they're doing, how effective their solutions are, and how long it takes them to do it. That's the basics of it across a wide range of professions. Now, there's no one-size-fits-all metric or formula you can just calculate based on objective facts for most of this, because the work is more varied than e.g. factory work, but it's also not impossible to make the comparison, if you actually understand the work reasonably and you use judgement.
In the case of this study, because the assignment of the comparison they were doing was random, then just measuring time to completion across a range of tasks is a perfectly reasonable metric, because there's nothing to really bias the outcome, just a lot of factors that add noise instead. But it is worth noting that the result is a very broad average, and there is likely a very complicated distribution of details underneath, which is much harder to measure.
AKA, be subjective! Which people are wary of, because what it brings is politics and tribalism.
Like I get that in SWE (like all other fields), managers have to make judgement calls and try to evaluate which reports contribute the most, but the GP post seemed surprised that this wasn't a solved problem by now, which just seems incomprehensible to me.
At the end of the road. Patient outcome and contentedness compared to others with similar indications. Patients seen and all that is that sort of short-term BS that you see everywhere that's giving metrics a bad name. It'd be like determining a mechanic's productivity by how many times he twisted a wrench.
Which would incentivise doctors to refuse to treat patients who are more ill, lest they risk their ratings go down.
Well I would first of all remark that this doesn’t seem and to be how it’s normally done as I’ve never been asked to rate my “contentedness” or similar with my medical care.
And where is the “end of the road?” Most medical interventions could be plausibly evaluated at all manner of different intervals.
Also, “similar indications” is doing a lot of work here. Patient outcomes are often influenced more by the individual than the doctor. By the time you bucket all the patients by age, diet, activity level, smoking status, alcohol intake, metabolic health, bmi, family history, etc…buckets are going to be pretty tiny. Clinics and hospitals aren’t that big, there won’t be anything to compare. If you only bucket the most obvious categories like age, you’ll have comparisons, but it will just be noise.
and how you would achieve it? "similar indications" would be coming from doctor that you are trying to rate
rating "contentedness" gets you doctors prescribing useless medications to keep patients happy
expert surgeons have often bad survival rates as they get complicated cases, and trying to rate how complicated cases are to compare two experts would be nightmare as bad as rating doctors - so you only replace one hard problem with another as hard problem
We have that shit like that too in SWE. Lines of code, github issues closed, features shipped, etc…
The hard thing is occupations where the quantity of effort is unrelated to the result due to the vast number of confounding factors.
https://secondthoughts.ai/p/ai-coding-slowdown
HN discussion: https://news.ycombinator.com/item?id=44526912
My entire life, I have written “ship” software. It’s been pretty easy to say what my “product” is.
But I have also worked at a fairly small scale, in very small teams (often, only me). I was paid to manage a team, but it was a fairly small team, with highly measurable output. Personally, I have been writing software as free, open-source stuff, and it was easy to measure.
Some time ago, someone posted a story about how most software engineers have hardly ever actually shipped anything. I can’t even imagine that. I would find that incredibly depressing.
It would also make productivity pretty hard to measure. If I spent six months, working on something that never made it out of the cręche, would that mean all my work was for nothing?
Also, really experienced engineers write a lot less code (that does a lot more). They may spend four hours, writing a highly efficient 20-line method, while a less-experienced engineer might write a passable 100-line method in a couple of hours. The experienced engineers’ work might be “one and done,” never needing revision, while the less-experienced engineer’s work is a slow bug farm (loaded with million-dollar security vulnerability tech debt), which means that the productivity is actually deferred, for the more experienced engineer. Their manager may like the less-experienced engineer's work, because they make a lot more noise, doing it, are "faster," and give MOAR LINES. The "down-the-road" tech debt is of no concern to the manager.
I worked for a company that held the engineer Accountable, even if the issue appears, two years after shipping. It encouraged engineers to do their homework, and each team had a dedicated testing section, to ensure that they didn't ship bugs.
When I ask ChatGPT (for example) for a code solution, I find that it’s usually quite “naive” (pretty prolix). I usually end up rewriting it. That doesn’t mean that’s a bad thing, though. It gives me a useful “starting point,” and can save me several hours of experimenting.
The usual counter-point is that if you (commonly) write code by experimenting, you are doing it wrong. Better think the problem through, and then write decent code (that you finally turn into great code). If the code that you start with is that as "naive" as you describe, in my experience it is nearly always better to throw it away (you can't make gold out of shit) and completely start over, i.e. think the problem through and then write decent code.
I find they often cause more trouble than they are worth, because they are completely wrong, and need to be “unlearned.”
"Experimenting" is a vital part of my process. I call it "Evolutionary Design,"[0] and it involves a lot of iteration. I have found that it's vital to UI[1], because I can almost never predict how UI will act, when actually presented to the user. The same goes for a lot of communication workflows. I have to "run it up the flagpole, and see who salutes." I almost always find that my theorized approach has issues, and I need to make changes. The old "Measure twice; cut once" approach to software development has caused me great trouble, over the years, and I have found that I need to adjust to new tools, and new contexts.
For example, right now, I am revamping one of my UI widgets[2]. It started as a minor tweak for iOS26, but I realized that it's a bit "long in the tooth," and that I can make it more robust, simple, and usable. I have been running the test harness all morning, seeing issues, and going back to the code, and tweaking.
[0] https://littlegreenviper.com/evolutionary-design-specificati...
I suppose it would be simpler to compare productivity for people working on standard, "normalized" tasks, but often every other task a programmer is assigned is something different to the previous one, and different developers get different tasks.
It's difficult to measure productivity based on real-world work, but we can create an artificial experiment: give N programmers the same M "normal", everyday tasks and observe whether those using AI tools complete them more quickly.
This is somewhat similar to athletic competitions — artificial in nature, yet widely accepted as a way to compare runners’ performance.
Focusing on the speed of outputting code or closing tickets is shortsighted from a software engineering standpoint.
Management is forced to rely on various metrics which are gamed or inaccurate.
The asymptote approached by software engineering is GDP=$0 because all problems are solved by maximally efficient automation. Never gonna happen, but progress on that path is a decent proxy for efficiency.
Too often the job is about introducing problems that are good for some company's bottom line, but that's the opposite of efficiency.
Some of the most productive devs don't get paid by the big corps who make use of their open source projects, hence the constant urging of corps and people to sponsor projects they make money via.
What about countries? In my Poland $25k would be an amazing salary for a senior while in USA fresh grads can earn $80k. Are they more productive?
... at the same time, given same seniority, job and location - I'd be willing to say it wouldn't be a bad heuristic.
10k PLN monthly would be a very good salary - and also perhaps a norm, as far as I know from experience and talking to peers.
I never understood how there can be a significant wage gap between these countries and Western Europe for jobs that are anyway done online.
I made great money running my own businesses, but the vast majority of the programming was by people I hired. I’m a decent talent, but that gave me the ability to hire better ones than me.
Changing jobs typically brings a higher salary than your previous job. Are you saying that I'm significantly more productive right after changing jobs than right before?
I recently moved from being employed by a company to do software development, to running my own software development company and doing consulting work for others. I can now put in significantly fewer hours, doing the same kind of work (sometimes even on the same projects that I worked on before), and make more money. Am I now significantly more productive? I don't feel more productive, I just learned to charge more for my time.
IMO, your suggestion falls on its own ridiculousness.
ON the other hand it makes no sense from some points of view. For example, if you get a pay rise that does not mean you are more productive.
i have to bill my clients and have documented around 3 weeks of development time saved by using LLMs to port other client systems to our system since December. on one hand this means we should probably update our cost estimates, but im not management so for the time ive decided to use the saved time to overdeliver on quality
eventually clients might get wise and not want to overdeliver on quality and we would charge less according to time saved by LLMs. despite a measured increase in "productivity" i would be generating less $ because my overall billable hour % decreases
hopefully overdelivering now reduces tech debt to reduce overhead and introduces new features which can increase our client pipeline to offset the eventual shift in what we charge our clients. thats about all the agency i can find in this situation