The team _will_ find out, and then instead of contributing to the success of the business in earnest, they’ll be doing stupid things like maximizing their changesets or racing for “easy” large changes like deleting a module.
It doesn’t matter whether those values do or do not correlate with reality (IMO, if they do, it is for relatively junior engineers only). If you give off the smell of measuring people like that, you will ruin any collaborative team environment and you risk never being able to recover that.
There’s a good chance you’ll chase away excellent engineers with this sort of low-effort metrics management too.
There are exceptions, but these are extremely rare.
Look for code created in collaboration by multiple engineers, with contributors or reviewers across multiple teams. This code is more likely to be an asset.
(Unless created by Google for external use or by xOoglers. Such code outside Google is a near guaranteed liability.)
Contrary to some other comments, I'd argue it's the people who ship lots of value, but very little code, who are rare to find in practice.
Looking for healthy collaboration is the key, this is where the magic usually happens. Counting lines of code is a poor proxy for this.
A single opinionated developer laying the foundation and core architecture is important for getting a coherent developer experience.
The other outcome if that developer is a good leader and interested in developing a community is that a community builds around the core foundation that was laid and begins to take a life of its own.
Look at React. Started by a single opinionated developer who laid the foundations and built a prototype. From there it's taken on a life of its own and evolved. It was not designed/built by committee from day 1.
Or python for that matter.
Counting them alone won't, but incentivizing people based on changesets will. IMO there are plenty of cases where counting LOC is useful to a manager. It's a discrete piece of data than can be combined with other data to make informed decisions as a manager.
For example, I have just written a Snowflake stored procedure that auto generates SQL code and then executes it. I am going to call this stored procedure from multiple places in my Python codebase. Now if my manager would look at the number of lines of code I would still write the stored procedure, however, rather than auto executing the SQL code, I would just copy and paste the auto generated code into my Python codebase increasing my LOC for that week.
Not everyone panders exclusively to optimizing grading metrics. I wouldn’t do that even if my manager looked at lines of code because I care about the quality of my work. I’m confident a lot of people share my sentiment.
If you are trying to use lines of code to measure productivity, the only reason is because you think some members of the team are not being productive. Sure, the good and productive team members won’t try to game the metrics… but the bad programmers, who you are trying to find, sure will.
The good programmers won’t game the system, but that doesn’t matter because they are going to be fine no matter how they are evaluated. So what is the point of measuring lines of code? You aren’t gaining any new information.
Bottom 10% get fired each year. What about now?
Another major confounder is that super deep investigations that fix major problems often have relatively trivial solutions at the end. But the effort to debug and fix them (that is largely hidden by conventional metrics) is enormous. The people working on these types of things are often your best devs (because you know they will actually get to the root cause), but if you are only looking at LOC you wouldn't know that.
It's just like shift times for workers in manual labor jobs. You can't tell which of the people who are meeting up 8 hours a day is most productive based on if they where there for 8 hours and 12 minutes or 8 hours and 25 minutes. But you sure as hell can easily see that if someone is showing up 4-5 hours on a nine-to-five job, then something is wrong. And if that person tries to "fake it" but showing up for 8 hours but spends half the time slacking off, but coworkers and managers will call him on the bullshit.
The idea the manager would call out individuals is laughable too. Most managers barely know how to code themselves. Give it another few decades before your manager looking through commits is a standard, you're more likely to see alert tools be prevalent before that.
You just need to create a shitload of bloated bullshit that somehow "works" (at least more or less).
That doesn't make you productive.
Actually that's even counterproductive as it creates technical debt. But by the bogus metric that would be positive.
I'm still to find the place where this is true in software development.
And when I then do actual work in between I lean towards the solution giving me more lines. Longer variable names, more line breaks, ...
There are so many ways to manipulate that, one I know it's a key metric. Sure, it's hard to go from weakest to top contributor that way, but good enough to get away from last spot who is being fired. At least till all play that game. By then the project goes towards a wall.
So yes, for a one off case to identify engineers one should talk to it is a somewhat useful metric, but I would look at more business relevant metrics like amount of closed bugs or similar first (while of course they can be manipulated easily as well ...)
- patents written / rewritten / meetings with lawyers / issued - mentoring of other engineers - PRs reviewed - talks given - meetings attended - RFCs / ADRs written - customer calls attended - quality of code (the original and best reason that LOC is bad.)
Et cetera, et cetera
If you are truly an engineering manager, unless all your direct reports are fairly junior, I sincerely hope you realize soon that LOC is a pretty terrible measure of engineer contribution.
I am old enough to remember when the industry finally came around to accepting LOC is a worthless measure of engineer productivity in the late 90s. I am not exactly shocked but dismayed that this conversation needs to be rehashed.
Number of tasks closed is a better measure than LOC and it is still a terrible measure.
Equivalent would be saying “our sales rep only said 2300 words today but that is ok, we can accept it he was visiting an old customer not cold calling like the juniors”
Usually LOC modified correlates with actual productivity but there are lots of exceptions. Like if someone copies an open-source library directly into the repo, commits a formatting change, is working on a critical section of code, is working in a different language, is writing new features while everyone else is fixing bugs, has been told they are going to be measured by LOC, etc.
The fact that you actually analyze your team members' contributions, know the codebase, and measure productivity in more ways than LOC is a good sign. The problem is when managers exclusively measure LOC to gauge productivity and make decisions. While LOC and productivity correlate more often than not, I bet you can spot at least some cases where the LOC is not an accurate measure.
Maybe for junior engineers and that’s it. I want my seniors and leads mentoring, doing code reviews, designing, communicating with the business about impact and the like, training, and yes, coding, but not as much as juniors.
Those seniors that maintain their output are only able to do that as long as they maintain their knowledge of the codebase. As soon as you lose control your output goes from 10x to 0.1x in the blink of an eye. At that point your options are to look like a fucking idiot because your juniors are out-doing you, or to optimise for blowing smoke up the manager's ass. Easiest way to do that is large sweeping rubbish architectural changes that look good on paper.
Some people will take a stand, but most just play the cards they're dealt. Every great engineer I know has had both 0.1x projects and 20-50x projects. Knowing why that's the case is more important for a senior/lead dev than coding ability IMO. Get the best out of others and all that...
As I’ve become more senior, I’ve found ways to push back on product requirements that add little value for the users, but tons of complexity in the code. If I don’t follow the requirements exactly, I can usually write ~10% of the code. But you need to learn how to do that with lots of experience.
It takes the same amount of time, because I need to dive in deeper to understand the problem and how it interacts with my features, lots of planning and meetings to change the scope, but the final output might only be 5-10% as much code if we can reuse things that already exist in the codebase.
—
For example, product team wanted something to be bold that we had set up to be within dynamic plaintext. It would’ve taken a rewrite of that entire UI component to be able to make some bits of text dynamically bold. But I was able to push back and say “what if we just rearrange it slightly so that the bold is unnecessary”, and the designers+business were ok with that too!
If I hadn’t pushed back, my more junior teammate would’ve written 2k+ lines of code to support that requirement. But instead it was probably like 10 lines to support the alternative formatting. That 10 lines took lots of time to understand the requirements, but was much better for code cleanliness and consistency in the end.
People will retort: "But, but, but... what about pushing junk?" Well, if you've decent code review system how will that get merged? Pushing junk is often a less of a problem provided you've a good engineering culture. The best teams I worked with were just plain fast. The worst ones had plenty of engineers hiding behind meetings, reviews, and bike-shedding.
[1]: https://medium.datadriveninvestor.com/when-overachievers-sel...
Depends what you mean by "productive". If you mean "appear to be busy" then yes. But by using this metric you disadvantage people who measure twice before cutting once, people who think deeply about a problem and find a simpler way to solve things.
In my experience, more junior people tend to spin their wheels a lot, and out of that spinning comes out a lot of code. Code that is often broken and then is is fixed again, and again. And such people appear to be more productive and busier, at a first glance.
Lines of code rewritten (changed N days after)
Lines of code removed
Lines of code that contributed to linting problems
Lines of code that contributed to security issues (as reported by the static analyser)
Average complexity measures
CI/CD build failure rates
Lines of code reviewed
Gitlab has plugins to generate reports like this, and if you have enough data (multi year preferably) you can use this as a springboard for deeper analysis of members of the team, I find all these metrics together (non taken as absolutes because there’s many assumptions built into each one) can give a good overview of individuals in a team
Build failure rates encourages frustration as long as you don't provide ways to run the exact tests in isolation locally, under the same rules as CI. Meanwhile, branch triggers and pre-push hooks trivialize this further.
Most of these metrics are strange given they are trivial to check pre-merge. Meanwhile, running these statistics pre-merge discourages small commits and sharing code prior to having ideal builds.
The moment this gets out is the moment devs will game your statistics and any benefit you gain goes out the door. And almost every one of these, even combined, is easy to game.
If they are important enough to measure, why not hard force them and focus on the remaining metrics instead?
And
“Use these as a springboard for deeper analysis of devs”
These are only used as high level metrics to guide further analysis, we are keenly aware of the faults of these metrics - but the alternative of no-metrics or worse, “gut feel”, is slow and riddled with more problems.
It doesn’t mean a dev is bad, maybe it means there’s some tooling failure, or organisational problem, or personal problem they need help with. These metrics give visibility into anomalies, and their imperfect nature doesn’t mean they are useless.
There’s also like 20 more metrics but I’m on my phone and couldn’t recall them all right now
Also I don’t know where all of this nonsense about “forcing” to run CI/CD before branch merges and “slap on the wrist” - all our devs can run the test suites locally in docker and inside their IDE, nobody is forced to do anything in the pipeline severs, and also nobody gets a slap on the wrist for CI/CD failures, actually I like to see larger changes with more failures because it means the dev is trying and I consider this normal behaviour, your incorrect assumptions about our test environment and how we interpret these metrics sound like you’re projecting.
i.e. you are not a competent manager by the usual understanding of «competent».
The metrics actually showed no surprises when I saw them, had I been asked to name the worst 5 devs, I would have named the same people that the metrics would have highlighted.
I raised this suggestion with management, but since we had a handful of devs who were drastically underperforming we decided to help them out first with increased mentorship, in order not to embarrass them in front of their colleagues.
I don’t know why you’re assuming the worst and malicious intent, it’s a great place to work, especially when we have a good vibe and all enjoy working together.
I think what contributes to this positive environment is not employing developers who instantly assume malicious intent, and aren’t paranoid about being discovered as incompetent. These people are usually the ones most opposed to metrics, and also are the ones who the metrics show are the worst performers. Instead of objectively assessing the facts they attack the messenger, full of emotion and vitriol because of their insecurities and inabilities.
There are many factors that go into firing a dev, if someone needs help we mentor them and grow them, but if someone has a bad attitude, assumes the worst and is full of negativity, then we are very quick to fire them.
Maybe your company is still a great place to work in spite of that. Or maybe it's only a great place for you to work, and your colleagues would feel differently if they knew about this. I can't say for sure because I don't have any further insight into your company. But that's also not the point about why it's unethical.
Nor is it having a 'bad attitude', assuming the worst, or being full of negativity to be opposed to the secret use of metrics in this way.
I see the metrics as a chance to provide visibility and support, you see it as a way to “manage” the devs. I see a chance to help, you see tyranny.
I see my viewing of my own metrics as a chance to introspect with an open mind and a willingness to see fault in myself and improve, you see an abuse of authority and a land grab to get ahead.
Then you say “I’m not assuming the worst, or full of negativity” when clearly you are, you just demonstrated it.
So like you, let me fall back on the ethics argument:
I think it’s unethical for a software manger not to take machine collected metrics and to manage on “gut feel” and intuition, just as all gut feel and intuition are not explained as rationale to all decision making (thus are also “secret metrics” - i.e. collected without transparent knowledge of all reasoning)
I think it’s unethical not to have secret metrics and inform every person of every thought and data point you have, so there aren’t any secret measures. (You better call your bank, your insurance company and loan companies and marketing companies who are all collecting secret metrics on you, in the sense their rationale, data points, and algorithms are not 100% open and are therefore secret)
I think it’s unethical to distribute reports that might demoralise some devs without helping them out first, especially if we know some of these reports might look unfair, and we are only using them as springboards for further investigation, like I have mentioned many times in this thread.
I think it’s unethical to complain about ethics and not address any of the points in the above discussion with objective facts.
But ethical things are ultimately subjective, so our conversation ends there, we will have to agree to disagree.
What company are you working for? Human "resources" department, I guess?
Looks like a place to strongly avoid because of very bad ethics and culture!
Your bank is scoring you by machine, so is your insurance company, and loan companies, are their algorithms 100% open and publicly available? No. Because they are secret metrics. Your bank knows more about you than you know, and so do the loan companies, I know, I used to work there. The datasets the loan companies have are insane, I could see if you are in financial distress or not, the value of your house, whether you were a “complainer” or not, your favourite music, whether you’re a foodie or not, and heaps, heaps more, I think 250 attributes on each person in the USA at one company.
What difference does it make if some metrics are collected by machine and not fuzzy, vague, human intuition? I think it’s an improvement.
Yes the metrics inform the seniors decisions partially like I said above, and yes they are taken with a grain of salt, like I said above, they are inputs for deeper investigation, that’s it.
Here’s an example that happened just the other day: I noticed heaps of CICD failures for one dev and I reached out to him. He told me he couldn’t easily run the test suite, so I spent an hour peer coding with him and showed him how to use VSCodes debugger.
1 day later he called me up and told me that “using the debugger was a life changer” and he was thrilled to be working so quickly, the faster feedback cycle from code->debug had restored his enthusiasm for work.
> Lines of code written
Less is better, right?
Because it's really easy to throw code at a problem.
OTOH it's hard to come up with an optimal minimal solution.
> Lines of code rewritten (changed N days after)
Means: Your devs commit unfinished, not thought out crap.
Or: Management is a failure because they change requirements all the time.
> Lines of code removed
More is better, right?
Because cleaning up your code constantly to keep it lean is key to future maintainability.
But could be also a symptom of someone with a NIH attitude.
Or, management is a failure because they don't know what they want.
> Lines of code that contributed to linting problems
If it's more than zero that's a symptom of slacking off and / or ignorance.
Your local tools should have showed you the linting problems already. Not handling them is at least outright lazy, and maybe even someone is trying to waste time by committing things that need another iteration later on.
Or, your linting rules are nonsense, and it's better to ignore them…
> Lines of code that contributed to security issues (as reported by the static analyser)
Quite similar to the previous.
Ignorance or cluelessness. Maybe even the worst of form of "I don't care", while waiting whether your crap passes CI.
In case someone doesn't know how to handle such warnings at all you have a dangerously uneducated person on the team…
Or, what is at least equally likely: Your snack-oil security scanners are trash. Producing a lot of the usual false positives (while of course not "seeing" any real issues).
> Average complexity measures
Completely useless on it's own.
Some code needs to be complex because the problem at hand is complex.
Also more or less most of this measures are nonsense. The complexity measure goes usually down when you "smear" an implementation all over the place. But most of the time it's more favorable to concentrate and encapsulate complex parts of your code. Hundred one-line methods that call each other are more complex than one method with hundred lines of code. But the usual complexity measure would love the hundred one-liners but barf at the hundred line method.
If there's someone who constantly writes "very complex" code (according to such measure done by some tool) this can mean that this person writes over-complicated code, as it could mean that this person is responsible for some complex parts of the code-base, or is consolidating and refactoring complexity that was scattered all across the place. Or maybe even something else.
> CI/CD build failure rates
This means people can't build and test their code locally…
Or someone is lazy and / or ignorant.
Or, equally possible, your CI/CD infra is shit. Or the people responsible for that are under-performers.
Or your hardware is somehow broken…
> Lines of code reviewed
How do you even measure this?
Just stamping "LGTM" everywhere quickly would let this measure look good. But is anything won by that? I guess the contrary is more likely.
Also someone who's constantly complaining about others code would look good here…
OTOH valuable and insightful code review is slow, takes a lot of time and effort but does not produces a lot of visible output. That's why I think it's disputable whether this can be measured even in a meaningful way.
There is only one valid way to asses whether someone is productive: You need to answer the question whether what this person is doing makes sense in the light of the stated goals.
But to answer this you need to look at the actual work / things produced by the person in question, and not on some straw man proxy measures. All measures can be gamed. But faking results is much harder (even still possible of course).
But not so good for comparing Joe and Mo. Or even saying “how is Joe doing”. Unless the reviewer is very skilled in navigating that data without bias. Most people I have worked for wont be and will ruin the employee relationship by making some comment about how “this could improve” when the employee things “ok this is BS but better go along with it, and maybe i’ll open CV.docx tonight and give it a spring clean”.
Bullshit is the biggest turn off for me in jobs and I will leave a job with too much (some is to be expected and a side effect of getting groups to work together). Team leaders tracking my time and commits closely to judge performance is this. How should they judge? Best way is get people to pair often. You will know what they are actually doing that way or can talk to someone who did. Any real laggers will be found naturally, and then find out why and fix (might be onboarding sux for example).
People do hate paring I know but it doesn’t have to be all the time. Even an hour a day of someone jumping on to help will spread boats loads of useful information.
As others point out though, the moment something becomes a metric it ceases to be useful. Check out the book "The Tyranny of Metrics" for a breakdown of it.