-2000 Lines of Code (2004)
folklore.org
folklore.org
-2000 Lines of Code - https://news.ycombinator.com/item?id=10734815 - Dec 2015 (131 comments)
-2000 lines of code - https://news.ycombinator.com/item?id=7516671 - April 2014 (139 comments)
-2000 Lines Of Code - https://news.ycombinator.com/item?id=4040082 - May 2012 (34 comments)
-2000 lines of code - https://news.ycombinator.com/item?id=1545452 - July 2010 (50 comments)
-2000 Lines Of Code - https://news.ycombinator.com/item?id=1114223 - Feb 2010 (39 comments)
-2000 Lines Of Code (metrics == bad) (1982) - https://news.ycombinator.com/item?id=1069066 - Jan 2010 (2 comments)
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu... (last one was 20 days ago, 1 point, 0 comments)
He started the lecture by analyzing how many pieces a machine could manufacture per day. Fair enough. He extended the model to measure different ratios of capacity. Makes sense.
Then he tried to extend the model to all machines, including humans. His example was: "How do you measure the capacity of a legal team?". I thought it was a trick question, so I answered (paraphrasing) "You can't answer that question the same way you answer for the machine. You can't give a single metric." He told me I was wrong and that the _right_ measure would be (total number of working hours/day).
I was tempted to try to convince him otherwise. The analogy was deeply flawed. He certainly measured the machines in (number of pieces / day) but measured the legal team in (hours/day). So, in analyzing a machine, you take into account its efficiency, but you don't do the same thing for humans.
I believe that is exactly the same thing that is going on in the post. Managers/Logistics/Economists are very susceptible to this kind of generalization pitfalls.
Edit: Given that this answer has generated some discussion I feel the need to expand on it. The legal team was not expected to sell their services "by the hour". In fact, any discussion about how their services were sold was shut down by the professor. From his point of view, the lawyers were machines and he was asking the question "how much can this machine produce?"
Yes, other students also suggested taking the number of billable hours/revenue into account, but that's not the answer the professor was looking for.
I don't criticize whether his answer is not technically right, but I feel it holds no real-world meaning. It was a purely academic question that leads nowhere instead of having a debate about how you measure the productivity of a group of human beings. And on top of that, his final answer was definitive and (from his point of view) was irrefutable.
Any such equation can only be represented by a single number ("capacity" or "productivity") if all variables are dependent (and therefore, there is only one independent variable in the equation).
So the assertion your professor is making is that the "capacity" of a team is always exactly dependent on hours worked per day, and any other proposed dimension of capacity (such as years experience, field of study, languages spoken, cases won, relationships with judges) are dependent on "hours worked per day". If he agrees any one of those variables affects capacity, but does not depend on "hours worked per day", then a single number can never reduce the dimensionality of the output (you need at minimum 2 numbers to represent two independent variables, you can never "collapse" the data).
> Any such equation can only be represented by a single number ("capacity" or "productivity") if all variables are dependent (and therefore, there is only one independent variable in the equation).
This doesn't seem right. You can have storage units with varied combinations of height, width, and depth, sure. But whether that matters depends on what you want to use them to store. An example of an approach that doesn't work would be storing unboxed fragile antique dollhouses. They have weird shapes, so you can't fill the floor area, and you can't stack them, so adding height to the storage unit doesn't add any capacity.
Except that of course you wouldn't just toss them into a garage and call it a day. (They'd break!) You'd keep them in boxes. Those pack and stack perfectly. Suddenly volume is what matters again, and increasing the width, length, or height of the unit by 10% will increase the amount you can store by about 10%.
This is even more obvious if you're storing water or oxygen. Fluids take the shape you give them. Your unit might have length, width, and height (though it really shouldn't... you want to store fluids in cylinders), but the only thing that matters for how much water you can put in there is volume.
However, air freight is a much more direct one. You have 2 largely independent measurements for weight and volume with either being the limiting metric for each load.
But I'm not saying that all multidimensional data can be losslessly reduced to a one-dimensional value. That would be crazy! the point of my comment is that it isn't true that -- as the parent comment asserted -- it is impossible to usefully report multidimensional data with a one-dimensional value. To the contrary, it is quite possible that the multidimensional data adds zero value over the one-dimensional summary.
Some multidimensional data can easily be losslessly reduced to a one-dimensional value. We can easily make a much stronger claim -- all one-dimensional values are "reductions" of other, multi-dimensional characterizations of the same data. But they're not all useless! The number of dimensions you use to describe data is an editorial choice, mostly unrelated to the raw facts.
You can't really assume a 100% packing rate along with increasing x dimension by y% meaning you can store y% more stuff
It's even truer if you're considering larger expansions; increasing the length of your warehouse by 200% will mean you can store 200% more stuff regardless of how awkward the original fit was.
1 - Underlying defects will absolutely sink your downstream production rate.
2 - If you only measure widgets per hour, the machine will make more, but smaller widgets (See Soviet Nail factory story - https://skeptics.stackexchange.com/questions/22375/did-a-sov... )
3 - Quality has a quantity all it's own. Many times, better widgets will improve efficiency many times over their own cost of production.
4 - Your professor was real dumb.
I'm pretty sure it's due to exactly that sort of hubris that business school folks have invited so much disdain. The domain in which you're operating simply cannot be dismissed as an inconsequential detail.
Tangentially, there is a subset of law firms that do operate as if total hours worked is the only thing that matters. Over the past decade or so, they've been rapidly losing ground to law firms that, by not thinking that way, manage to do a better job of producing the kinds of output that clients actually want.
So, the question "How do you measure the capacity of a legal team?" (note it says capacity), makes sense. It's the answer I disagree with.
What I think that a lot of management type folks fail to realize, though, is that both the quality of knowledge workers' output and the rate at which they produce it tends to drop precipitously when they are tired. I wouldn't be at all surprised if a lawyer who works 35 hour weeks can get more done in a given calendar period than one who works 90 hour weeks. Big name law firms, though, bill by the hour, and, even if they share this conviction, they know that their clients went to business school, and have therefore been trained not to understand it.
As for my personal opinion, I haven't reflected on it too much, but I think capacity implies a quantitative (edit: measurable may be a better word?) output, but not necessarily fungible.
For lawyers, most of them sell hours. The more hours they bill, the more productive they are.
Most businesses who hire programmers do not make their money by billing programmer hours. So that metric wouldn't work. Lines of code seems reasonable until you think it through. Honestly, I don't know that anyone has come up with a good solution for measuring programmer productivity.
But lawyers? They're in the business of selling time in 15 minute increments. Their productivity is simple to measure in this respect.
Surely they care about their underlying activities and not just number of legal hours worked, right?
I have to believe there is a legal team somewhere in the world that is measured on productivity beyond just "number of billable hours" generated.
I think the problem is the word "productivity". When we say that word, we're implicitly suggested there's a simple integer or decimal that can capture whether a person's wages are money well spent or not. For most professions, programming included, I am highly skeptical of the existence or even potential for such a number.
> But lawyers? They're in the business of selling time in 15 minute increments. Their productivity is simple to measure in this respect.
I don't think all lawyers are in that business. There are plenty of in-house counsel that aren't in that business. Just as there are plenty of engineers and software developers that _are_ in the business of selling time in 15 minute increments.
I just don't think the productivity question actually breaks down along professional lines, but rather on business model lines (which, again, I think we're in agreement about your main point)
I also agree it doesn't break down along professional lines. Just used lawyers as an example, but I shoulda been clearer about my intent.
I would not call it moot, but opportunity cost - and opportunity cost is very hard to measure. Technical debt is similar.
If you have current lawsuites to handle and avoid cost, their hours are better spend doing that than dishes. If you have future law suites to avoid... that get's even more tricky.
If the hours aren't tied to either the cost or the price, then I don't know how they can be tied to productivity in an economic sense.
Sure, but at some point, you can pick a definition that is so far removed from what was intended, that this exercise is utterly meaningless.
You could say that anytime there's any lack of clarity about what is meant by any given term,
With "productivity", you could reasonably mean any number of things.
It's not like someone said "pizza" and I said "that depends on what you mean by pizza". You could say that (is a calzone a pizza?), but it wouldn't be reasonable to do so.
In the case of productivity, I think it's reasonable to clarify what is meant.
P.S. Was your use of "utterly meaningless" an intentional pun?
Well, the only people who could meaningfully search for such solution - programmers - have all the incentive in the world not to find it. Not very surprising they didn't find it yet (and won't ever.)
As an Engineer, I sell my time by the hour too. No different than the lawyer. Yet I try to finish things efficiently. Huh.
As in you literally bill for hours, and the more hours you work, the more you get paid?
Most programmers that I know (which is obviously not a great metric) either get paid a salary (which is divorced from actual hours worked) or they get paid by the hour but have no say over how many hours they will. In both cases, time is independent from productivity. Therefor, there's no harm (and really only benefits) to coding efficiently.
But if you a) control how much time you work (like lawyers do, to an extent), and b) get paid for your time, then yes the incentives are setup to encourage you to be inefficient. Completely agreed.
If you bill by the hour, there's a reasonable (but not necessarily correct or optimal) case to be made for measuring your productivity in terms of billable hours.
But it makes zero sense to do so if your hours are disconnected from the economic activity that results from your work. In such cases, it would be completely arbitrary to measure hours and call that a measure of productivity. Just as arbitrary as lines of code.
that's the way I do it.
>then yes the incentives are setup to encourage you to be inefficient.
I'm pretty sure if I was inefficient I would lose the job and thus not make as much money as I would otherwise.
The software engineering profession is riff with wasted work: rewrites, new bullshit services and tech, insanely complex clustering and cloud deployments, etc.
Most lawyers I know have many clients and are swamped with work. They have little incentive to "stir things up". It's similar with accountants and plumbers in my city. They aren't trying to make more work for themselves because they already have their hands full.
But in regional markets where supply isn't so constrained relative to demand then, sure, there's an incentive to make-work once you've wrangled a client, just as with any other profession.
If the team is doing work for the firm, but you don't want to complicate the model, you can stick a labor-enhancing constant (to allow for heterogeneity between workers) and use "work" as a unit. Sure this model is wrong, but all models are wrong. We're just trying to create some useful ones.
Doesn't this mean that by your professor's metric they are becoming less productive?
More cynically, you could measure a racer by the amount of revenue generated by sponsorships, ad placement on the yacht hull, endorsement fees/kickbacks, etc.. If you have two equally competitive racers, but one is more mediagenic, perhaps that one has higher "productivity"? If a racer often loses, but does so in engaging, nailbiting ways that create a following, perhaps that one is "productive"? A wrestling "heel" may lose their bouts but be a successful character, say.
Yup, they could start sabotaging their competition or bribing judges to disqualify other competitors, etc.
I always argued that the value we provided may not be realized for years, when some clause we put in a contract prevented us from getting sued. We may not even know it had happened! Ultimately, for non-litigation legal work, your value as a lawyer is often in preventing bad things from happening. How do you measure that?
So yeah, your professor compared apples and oranges.
1. https://people.wou.edu/~shawd/mediocristan--extremistan.html
2. https://kmci.org/alllifeisproblemsolving/archives/black-swan...
Thank him for his "wisdom" but make sure you do it in a convincing way. Finish the class, get your A, move on with your life. You don't need him to acknowledge he's wrong, you just need him to give you good marks.
(There are obviously some “jobs” a lawyer does that tend be charged at a fixed rate such as conveyancing, which you would want to optimise for throughout)
I have updated the post to expand on your observation.
If you want everybody to guess what you are thinking, become a quizmaster.
Lawyers are often billed by hour, but that's not what they are paid for. You will quickly lose customers if you optimize for that metric. Instead, if you figure out how to spend less hours for the same results, you may charge even more per hour, because you are saving your customers' time, plus handle more customers over a given time frame.
But as with lawyers, programmers, managers and all other non-machine-like workers produce "customer happiness" that's not easily measured, other than at the level of overall competitiveness of the firm.
> Edit: Given that this answer has generated some discussion I feel the need to expand on it. The legal team was not expected to sell their services "by the hour". In fact, any discussion about how their services were sold was shut down by the professor. From his point of view, the lawyers were machines and he was asking the question "how much can this machine produce?"
Pricing the value of in-house counsel is an interesting problem in itself (because they typically don't bill by the hour). One could use an insurance policy pricing model (the worst that could happen is a very costly loss in a lawsuit) to determine what's the counsel protecting the company from.
My local university. If you want the specific school, my handle is associated with my real-world identity. It won't be hard to figure out.
> Pricing the value of in-house counsel is an interesting problem in itself (because they typically don't bill by the hour).
Absolutely!
> One could use an insurance policy pricing model (the worst that could happen is a very costly loss in a lawsuit) to determine what's the counsel protecting the company from.
Also true. The lack of cost analysis in my classes worries me very much.
Time to transfer somewhere else?
Generalizations or rules of thumb can actually outperform complex quantitative approaches to decision making. Look at the 1/N heuristic for portfolio management for example.
Anyhow just my opinion - something for your curious mind to consider !
You could to say production line workers using piece rate system but a a legal team consists of many different people with different skills and who perform many different tasks.
Classes taught? Students taught? Students getting degrees (undergrad/Masters/PhDs)? Research grants won? Nobel prizes won?
A legal team generates income by billable hours (value $y) worked per day.
working hours = input
Would the professor be okay to pay me for the time I spend reading web? I mean, from his perspective, it is the time I spend that is important, not what I do.
This is necessary, because humans have very limited capacity to understand world around them and otherwise it would not be possible to make informed decisions, as gathering all relevant information would necessarily take practically infinite amount of time.
From that point of view I understand people like your professor is mostly result of bias also called Dunning Kruger effect. This is basically lack of education in a given area. You need at least some knowledge in an area to be able to appreciate complexity and unknown unknowns.
If you don't want to be that guy, the best medicine is first to learn to be self aware, second to be aware of various biases (including Dunning Kruger effect) that you are subject to and third to get some knowledge/experience in an area you are trying to make decisions in.
My similar story was removing a 10,000 line module that built hundreds of different packets for sending over a mailbox to a wifi module. Each method was almost identical, with the exception of building a slightly different header.
I replaced it with a template that given the structure, built a packet to send it. Less than 1 page of code.
It was like a computer science geek gone mad had figured, "I'll use every decomposition technique I ever heard of, and invent some more, so this is the most computer-sciencey source base in the world".
Ultimately I rewrote it in 12 C++ base classes and a derived object for each radio card I had to work with.
The number of times I've seen:
if (my_var == 'some value' || true == false) { ... }
Why do people do this instead of commenting out the code?!The thing I really can't understand is why you would compare a boolean condition to `true` instead of checking the condition itself (in other words, writing `false` instead of `true == false`). Also, let me suggest to use `&&`, not `||`. And while we're at that, I would actually write:
if (false && my_var == 'some value') { ... }
just in case operator `==` can have side effects or relevant computational cost in your language. if (my_var == 'some value') { ... }
which is not the same as commenting it out entirely.Looking at history via "blame" is useful to see why a bug was introduced (was it fixing another bug, and if so, is your fix going to break that bug again?), and how long it's existed for.
Leaving old code commented out doesn't help either with of those things. Unless maybe it's accompanied by lots of comments and date stamps, in which case you've not only re-invented a very crappy, half-baked version control system, but also made the code hard to read and work with.
Git blame also doesn't help when the history gets truncated for performance reasons.
But overall, if commented code ends up in version control, it probably could have been removed.
if (my_var == 'some value' || true == false) { ... } //if (my_var == 'some value') { ... }
if (my_var == 'some value' || true == true) { ... } //if (true) { ... }
if (my_var == 'some value' && true == false) { ... } //if (false) { ... }
That said... I would never use the comparisons `true == true` and `true == false`... that's just sillyRight? Wouldn't it be more readable to use just 'true' or 'false' instead of the comparisons, or is that not a feature in some languages? I don't understand what might be gained from the extra comparison, besides confusion.
if (my_conditional || true) { etc }
if (my_var == 'some value') { ... }But that's arguing syntax instead of principle, I think generally "less code" means "fewer expressions to evaluate"
Of course, my manager is telling me to keep adding features to the old codebase. Sigh.
Most complex processes can't be reduced to a handful of simple variables - it's oversimplification at its worst. The best you can do is use metrics for a jumping-off point for where something /might/ be going wrong and thus start engaging with actual humans (or reading code/logs/some other source of feedback). Too often I've had to deal with management who go straight from metrics to decisions and end up making bad decisions (or wasting everyone's time shuffling paper to generate good looking metrics).
You then state what seems to be the mainstream view on HN. Certainly I don't see it as controversial, just kind of obvious
Goodhart's law rephrased by Marilyn Strathern: "When a measure becomes a target, it ceases to be a good measure" https://en.wikipedia.org/wiki/Goodhart%27s_law
Campbell's law: "The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor" https://en.wikipedia.org/wiki/Campbell%27s_law
The problem is that number of lines of code is a good metric only if it is decreased without decreasing code quality.
The smartest people who are stamping parts all day were the people who Ford promoted for higher positions to make the work more efficient. He wrote that most people were happy with the repeated work, but there were a few who were better as leaders or engineers.
Tesla’s growth curve is actually very similar to what Ford’s was at the start.
It's managers treating software development like an assembly line that leads to waterfall, management Taylorism and other proven-to-fail concepts. You can optimise for innovation or you can optimise the speed of a repeatable assembly line but one business unit can't do both in the same framework because optimising speed requires reducing process flexibility and innovation requires increasing it.
To tell you the truth I think waterfall model was not about getting the best manufacturing, but about the leaders not taking any risks and saving their own jobs.
SpaceX is clearly an innovative company and I'm sure they're not using a Ford-style assembly line because that would make no sense for a quality-over-quantity product like a rocket.
By assembly line I specifically mean a Fordian assembly line where units move between stations manned by specialists in a single step of the process.
That was the result of lots of innovation that Ford did. And then all the car companies stopped innovating on it.
For example Ford started to use electric motors for each machine separately instead of having 1 big motor that tried to power all machines. He sped up the assembly line by 10x at least and measured all operarions carefully.
The assembly line you are talking about is the last process set in stone for 100 years instead of innovating further.
If you point was that companies are generally bad at doing this, or that they often measure the wrong things, or that the process can be abused, or that you should not attempt to measure something beyond a certain level of precision, then I'd agree with you. But to write the entire process off as useless is just as unproductive as the problematic situation you're criticizing.
My point is that an obsession with empiricism can make you think that only #1 is valid evidence and thus use it for qualitative analysis where it should not be used.
Only using metrics for feedback is giving yourself tunnel vision.
Right now our metric is basically - talk to the developer and try and see if he's BSing you and goofing off, that's super subjective and very vulnerable to personal biases, but, it is a metric - it's just not an objective metric.
I don't know what it is - I've never seen evidence of a good one out there - but I don't begrudge managers trying to find new objective measures for productivity. I'd be quite excited to see one myself.
The idea that only truths expressable in abstract equations are objective and thus true is exactly the kind of false belief that gets us in trouble.
> Right now our metric is basically - talk to the developer and try and see if he's BSing you and goofing off, that's super subjective and very vulnerable to personal biases, but, it is a metric - it's just not an objective metric.
That isn't a metric. Metric, having the same root as metre, is about measuring. What you're talking about there is a heuristic, and they're much more effective for tracking qualitative issues.
How would you, as a leader, keep track of how your organization is running?
So in your example, you just have to rely on the judgement of all your professional project leaders and architects and what they tell you.
The specific approach to metrics I was referencing as being better is known in cybernetics as an algedonic alert. It doesn't seek or claim to provide information, it only rings the bell of "investigate this area", like a CloudWatch alert for your organisation.
Using metrics to make decisions is the mistake in my mind.
In my work, we shifted to online project management tools (without training, dare I say), which is just additional work on top of actually getting things done (and balancing with increased household maintenance, which nobody talks about).
Worst, we had also wasted meeting hours (everyone's time) just defining how to define our progress.
> ... wrote in the number: -2000.
> I'm not sure how the managers reacted to that, but I do know that after a couple more weeks, they stopped asking Bill to fill out the form, and he gladly complied.
Bill — but what about the rest of the team? The devil’s answer: They were expected to keep supplying the number, because line management was forwarding the stats up, having previously “sold” upper management on their value. And to admit error on such a fundamental is career-threatening.
Measuring programming progress by lines of code is like measuring aircraft building progress by weight. - Bill Gates
My personal point of view is that: every line of code you write is a liability. Code is not an asset; a solved problem is.Edit: To clarify, I'm definitely not encouraging writing "clever" short code. Always strive to write clear code. You write it once, but it will be read (and potentially changed) many, many times.
Gates's, on the other hand, accurately captures the reality: while, all else being equal, a lighter-weight design is preferable to a heavier one, it takes some skill and effort to actually produce the lighter design. Which leaves open the possibility that doing so may not actually be worth the effort.
In short: Make program logic similar to business logic.
Example: "Zero code solutions". Hell, we even had that back in the old days with Java frameworks doing a ton of configuration via XML files. Sure, it's not "code", but it's basically the same thing and still a liability. Especially if there's a very steep learning curve to all of the features/configurations that you'll never use.
A more subtle problem with abusing this idea is the over-dependency on third party code. Sure, it might look like you haven't written any code when you do `npm install foo`, but really, you've just placed a bunch of trust in some other person. Do you vet your dependencies' authors the same way you vet potential employees at your company?
There's a fine line and it's an art form figuring out when to bring in outside code or to NIH it. IMO, of course.
Anyway, it's interesting evaluating those potential solutions, because certain things like going to a daemonless, rootless, bind-less build based on podman/buildah is a no-brainer, but the next frontier beyond that has a bunch of tools like ocibuilder, cekit, ansible-bender, etc which want to establish various ways of declaratively expressing an image definition, and although the intent is good there, it's absolutely not worth getting sucked into long-term dependence on pre-1.0 tools with single-digit number of contributors and an uncertain maintenance future.
I utterly reject this assertion.
I'm wondering if there's a quality term for a "unit of complexity" in code? Like when a single line has 4 or 5 "ideas" in it.
and this one https://marketplace.visualstudio.com/items?itemName=Stepsize...
These two haven't tried before but doing it now:
https://marketplace.visualstudio.com/items?itemName=selcuk-u...
https://marketplace.visualstudio.com/items?itemName=TomiTurt...
switch (foo) {
case KIND_1: return parseObjectOfKind1();
case KIND_2: return parseObjectOfKind2();
case KIND_3: return parseObjectOfKind3();
...
case KIND_15: return parseObjectOfKind5();
case KIND_16: return parseObjectOfKind16();
}
There are 16 paths and yet this code is easy to follow. There is no substitute for human judgement (code reviews).[1] https://en.wikipedia.org/wiki/Cyclomatic_complexity#Correlat...
I don't like this type of code for exactly this reason.
switch(HTTP_METHOD) {
case PUT: processPut(...);
case POST: processPost(...);
case GET: processDelete(...); // wtf
case DELETE: processGet(...); // mate?
}
To the reader that appears to be an error even if it is precisely the thing you want to happen.I think everyone has coded this way out of expedience and some of us have eventually messed it up too. But you could, for example, use a macro in C to only list the number once. It might not be a trade-off worth making though.
Personally, I've used reflection to do this kind of mapping of enums to functions in a language that supports reflection.
Turning "box" into "parseBox" using string operations and then using reflection to call the right method is an approach I'd consider in Ruby, but not in Java. It breaks IDE features like "Find Usages" and static analysis, and the code to dynamically invoke a Java method is more annoying to review than the boring repetition in my posted snippet.
I think it's pretty clear that the number is a placeholder for something reasonable, e.g. making an association between two distinct sets of concepts. You'll still be vulnerable to copy-paste or typo issues.
> Personally, I've used reflection
Now you have two problems (and still have to maintain an association between two sets of concepts).
In my own code, I have used the enum name to map to a set of functions related to that enum value. The association is implicit in the name of the enum value and the name of the functions. No way to mess that up like this.
If it was:
case: KIND_BOX: return parseObjectOfKindBox();
It's no difference. Still repeating "Box".In more dynamic languages, you could probably use introspection.
Lastly, this does not alleviate the association issue, but I prefer the alternative of declaring the associations in an array/map somewhere, and using map[enum_value]() instead of the switch.
A map can be a good solution though, particularly if it's something like mapping enum values to strings that is constructed via EnumClass.values().map... or something so that you know the map is a total function.
It is the same with the switch, at least in C/C++. You either have a default, or must list all the possible cases and still return something "at the bottom".
In C, you could use a macro in this call to make sure that the name/number is only specified once.
I think the real issue is types aren't included in the example. I work in much higher-level languages, but if you are passing strongly typed objects thru and your switch is derived on this typing, it's probably going to be illegal in your type system to return certain results.
If your type system doesn't validate your results, then you'll be prone to the class of error you are discussing. Maybe that's common in C.
As for correlation to number of defects, another important factor is maintainability (velocity of new features in growing programs, keeping up with infrastructure churn in mature programs). If reducing perceived complexity improves maintainability, and if measuring cyclomatic complexity despite its flaws helps reduce perceived complexity, then it's still a useful metric.
And I think that is a problem. switch()ing over an enum will give you a warning about unhandled cases in some languages. But if you start building your own lookup tables, you will be just as prone to typos like mine, except with more code, and less static analysis.
Yet the cyclomatic complexity drops from 16 to 1. I don't think that's a healthy incentive.
I think it's easy to come up with an algorithm that's better than a coin flip or counting LOC. But to be useful, the algorithm has to be on par with the judgement of the developers who check in the code and review it. Otherwise the false positives will take time away from other bug-reducing activities like code reviews and writing tests.
Of course I don't have data to prove it, but I'm convinced that every complexity checker I've encountered in the last 15 years has been a (small) net negative for the project. Maybe machine learning will improve the situation.
Not that I'm trying to sell "ideas". I don't even know. But it's this very loose concept that floats around my mind when writing code. How many ideas are there in a stanza? Ahh too many. I should break out one of the larger ideas to happen before.
But it's only a mathematical construction and is uncomputable, just like the halting problem. In real life, for some applications a good compression algorithm like LZMA is sufficient to approximate it. But I'm not sure if it's suitable for measuring computer programs - it would still have a strong correlation to the number of lines of code.
https://www.sonarsource.com/docs/CognitiveComplexity.pdf
>Switches
>A `switch` and all its cases combined incurs a single structural increment.
>Under Cyclomatic Complexity, a switch is treated as an analog to an `if-else if` chain. That is, each `case` in the `switch` causes an increment because it causes a branch in the mathematical model of the control flow.
>But from a maintainer’s point of view, a switch - which compares a single variable to an explicitly named set of literal values - is much easier to understand than an `if-else if` chain because the latter may make any number of comparisons, using any number of variables and values.
>In short, an `if-else if` chain must be read carefully, while a `switch` can often be taken in at a glance.
I disagree with them on assuming method calls being free in terms of complexity. Too much abstraction makes it difficult to follow. I've heard of this being called lasagna code since it has tons of layers (unnecessary layers).
Maybe the complexity introduced by overabstraction requires other tools to analyze? It's tricky to look at it via cylcomatic complexity or cognitive complexity since it is non-local by nature.
Great! Now we can all begin to write branchless code. Make branch prediction miss a thing of the past!
I wasted a lot of brainpower cutting down CC in my code doing terrible things like storing function names as strings in hashmaps, i.e.,
String fn = codepaths.get(if_statement_value);
obj.getClass().getDeclaredMethod(fn).invoke(param,...);
Because that could replace several 'if' checks, since function calls weren't branches. Of course, no exception handling, because you could save branches by using throws Exception.To this day, I wonder if Cyclomatic complexity obsession helped make "enterprise Java" what it is today.
Minimizing cyclomatic complexity might actually be a reasonable approach, if the complexity were accurately measured without these blind spots. For example, any indirect calls should be attributed a complexity based on the number of possible targets, and any function that can throw an exception should be counted as having at least two possible return paths (continue normally or branch to handler / re-throw exception) at each location where it's called.
As with any metric, it's something to take as useful "advice" and becomes a terrible burden if you are optimizing for/against it directly. (Such as if you are incentivized by a manager to play a game of hitting certain metric targets.)
It's also interesting to note that a complete and accurate complexity metric for software is likely impossible due to it being an extension/corollary of the Halting Problem.
As for complexity metrics, this is a contentious topic. It's hard to quantify "ideas" or "concepts". I think this is the part where technical design reviews and good code reviews help keep in check.
Just like debt, there is a right amount to have: too much and you are incapacitated because you owe more than you can produce, too little and you don't have enough runway to produce what you want. Note that I'm talking about from a company's point of view.
The problem you might be avoiding is change.
Things that are bad should probably be fixed, though.
Very true... but also false in a way.
Here's a little bit of functionality that has to be in there to implement the way we're solving the problem. That code is a liability; it would be better to solve the problem in a way where we don't need this bit of code.
But given that we're solving it this way, let's say that this little bit of functionality may be implemented in one line, which is almost unreadable, or in five very clear lines. That one line is far more of a liability than the five lines are.
Easy to read code with fewer decisions should be the goal of a code minimalist.
Now I am a dependency minimalist (as much as is practical, it's a continuous gradient trade-off and naturally YMMV) more than I am a pure code-written-here minimalist.
I'll happily double my SLOC for most small apps if it means my app can be stdlib only.
E.g. we don't generally count (g)libc because a C-programmer should be familiar with the C API. We don't count Linux syscalls for the same reason, generally. But we might want to keep in mind that many APIs have dark corners that few users of the API are aware of, and so we may make exceptions.
But the more obscure the API, the more important it is to count the complexity of the dependency as well.
Both because it increases the amount of code your team must learn to truly understand the system, and because if it is obscure it also reduces the chance that you can hire someone to compartmentalise the need for understanding that dependency as well.
You may not be able to count lines of win32 code, and its awfully hard to make a patch, but if it's broken and you depend on it, your product is broken.
There's also a multiplier that should be attached though. Most products don't have developer time or skills to write their own OS, so there's value in using someone else's even if it's probably more code than a custom built one that only satisfies the needs of the product.
We can proxy a little but the core problem is that the function that spits out your metrics for those actually has a hidden parameter of the audience you are writing for and the purpose it is for.
So when the audience is highly familiar with Linux (the kernel and platform) idioms, you could choose an exotic microkernel with far fewer SLOC and actually have true lower score.
Of course that's pointlessly edge case, but the natural simple easier to understand version of that is just using a different utility library instead of the one currently used in the codebase. This one could be smaller by far and still be worse in truth.
Can still be fine for opensource/hobby work, anything professional needs better integration with the individual platform native UI apis.
Which is one of the reasons Electron became so popular — nobody has any expectations from a webapp UI, yet they still look better than Qt/GTK/wx on average...
So my opinion is if you want an app that looks perfect in screenshots or your customers are liking the buttonsand complain if the corners are not round enough then you should use the native looking stuff, but if your customers do their job using this app and every minute lost because of bugs. missing feature or bad UX then the look of the toolkit is not the issue, focus on what the customer is doing , see where you can improve his work and you will have happy customers.
As you add code, the best structure for that code changes and you want to refactor. I'm not just talking here about pulling some shared code into a new function, I'm talking about moving responsibilities between modules, changing which data lives in which data structures etc. These changes are the key to ensuring your code stays maintainable, and makes sense. Every unit test you add 'pins' the boundary of your module (or class or whatever is appropriate to your language). If you have lots of tests with repeated code, it can take 5 times as long to fix the tests as it can to make the actual refactors. This either means that refactors are painful which usually means that people don't do them as readily (because the subconscious cost-benefit analysis is shifted).
If - on the other hand - you treat your test suite as a bit of software to be designed and maintained like any other, then you improve this situation. Multiple tests hitting the same interface are probably doing it through a common helper function that you can adjust in one place, rather than in 20 tests. Your 'fixtures' live in one place that can be updated and are reused in multiple places. This usually means that your test suite helps more with the transition too - you get more confidence you've refactored correctly.
The other part of this problem (which is maybe more controversial?) is that I try not to rely too much on lots of unit tests, and lean more on testing sets of modules together. These tests prove that modules interact with each other correctly (which unit tests do not), and are also changed less when you refactor (and give confidence you didn't break anything when you refactor).
I guess my point was not that you never DRY in tests, just that you should be very picky about when to DRY, more so than in code, and that is necessarily in opposition to the advice in OP.
A test suite with a lot of factored-out common bits makes the tests harder to understand. It's similar to the worked examples in a math textbook. If half a dozen similar examples factored out all the common bits (a la "now go do sub-example 3.3 and come back here", and so on), they would be harder to understand than repeating the similar steps each time. They would also start to use up the brain's capacity for abstraction, which is needed for understanding the math that the exercises illustrate.
These are two different cognitive styles: the top-down abstract approach of definitions and proofs, and the bottom-up concrete approach of examples and specific data. The brain handles these differently and they complement one another nicely as long as you keep them distinct. Most of us secretly 'really' learn the abstractions via the examples. Something clicks in your head as you grok each example, which gives you a mental model for 'free', which then allows you to understand the abstract description as you read it. Good tests do something like this for complex software.
Years ago when I used to consult for software teams, I would sometimes see test systems that had been abstracted into monstrosities that were as complicated as the production systems they were trying to test, and even harder to understand, because they weren't the focus of anybody's main attention. No one really cares about it, and customers don't depend on it working, so it becomes a twilight zone. Bugs in such test layers were hard to track down because no one was fresh on how they worked. Sometimes it would turn out that the production system wasn't even being tested—only the magic in the monster middle layer.
An example would be factory code to initialize objects for testing, which gradually turns into a complex network of different sorts of factory routines, each of which contribute some bit and not others. Then one day there's a problem because object A needs something from both factory B and factory C, but other bits aren't compatible, so let's make a stub bit instead and pass that in... All of this builds up ad hoc into one of those AI-generated paintings that look sort of like reality but also like a nightmare or a bad trip. The solution in such cases was to gradually dissolve the middle layer by making the tests as 'naked' as possible, and the best technique we had for that was to shamelessly duplicate whatever data and even code we needed to into each concrete test. But the same technique would be disastrous in the production system.
Why we keep promoting these people into positions of management is beyond me.
edit: here's one that I find to be an interesting character study: https://www.folklore.org/StoryView.py?story=Round_Rects_Are_...
Generally by the time it'd reach us, it was something requiring more in depth troubleshooting.
They introduced a metric to measure ticket performance. The rough idea was "faster it's resolved, the better" (reasonable measure, if you're also tracking customer satisfaction), combined with "fewer interactions with customer the better" which was an absolutely stupid way to measure performance.
About a month after it came out, we were getting chewed out for our "conversion score" being low. Too many interactions with customers, and tickets taking a while to handle. No shit, we're the top tier of support. If it got to us it was bound to take time to resolve, and almost certainly involved a lot of customer interaction.
One of the engineers in the team managed to dig up how to get a "conversion" rate report up for any support engineer, though not the code that generated the figures, and very quickly realised that the way to get 100% conversion rate was just to resolve and immediately re-open the ticket as soon as you picked it up. We all promptly started doing that, and they stopped chewing us out.
If you incentivise the wrong behaviour, you're going to get results you likely don't want.
I wonder if we should name the companies who do this, or if it is fighting a losing battle? In the end, some management just wants to look at charts.
Edit to add: This happens outside of programming as well. I know a guy who worked at AT&T as a DSL Installation & Repair tech. They had such a focus on how long techs spent on a given job (less time was encouraged, more time was penalized) that a lot of his co-workers would go to the DSLAM and snip a wire so that they would be called out the next day to fix the problem. He pushed back so heavily on the poor incentive of "getting your numbers up" that he eventually got written up for insubordination. We need a larger eye-roll emoji.
I've never done this myself, but I've seen developers do similar things rather than fight management. Measuring and rewarding the right things is important.
According to Wikiquote there is no primary source to show he really said that. Nevertheless, Steve Ballmer did,
> "In IBM there's a religion in software that says you have to count K-LOCs, and a K-LOC is a thousand line of code. How big a project is it? Oh, it's sort of a 10K-LOC project. This is a 20K-LOCer. And this is 5OK-LOCs. And IBM wanted to sort of make it the religion about how we got paid. How much money we made off OS 2, how much they did. How many K-LOCs did you do? And we kept trying to convince them - hey, if we have - a developer's got a good idea and he can get something done in 4K-LOCs instead of 20K-LOCs, should we make less money? Because he's made something smaller and faster, less KLOC. K-LOCs, K-LOCs, that's the methodology. Ugh anyway, that always makes my back just crinkle up at the thought of the whole thing."
Not really a fan of either of them, but we can all agree on the quote.
See CIA (OSS) manual for details.
https://www.openculture.com/2015/12/simple-sabotage-field-ma...
Whenever I hear people in the automotive industry boast about the complexity and lines of code in vehicles I weep and shake my head.
Reference: https://youtu.be/YAtLTLiqNwg?t=953
The change, annualized, was somewhere around $12M in profit.
I'm still pretty proud of making -$3M/LOC!
Ok, sure, sounds easy. Then I opened the script... 9000 lines. After reading and understanding what it was supposed to do, I rewrote it in 1500 lines (still basic SunOS unix shell code), and with reasonable use of internal data structures for caching so only the first menu visit required a time hit. Beyond that, it was 1 second for menu selections. To say the system engineers were pleased would be an understatement.
My manager was pleased but also displeased, because he was the author of the 9000 line monstrosity.
I contributed an inliner to a language about ten years ago. Inlining is a problem that might seem easy at first, but for me it was like trying to restrain a rabid dog on a leash. I was pretty damn pleased with the end results and it served the language implentation well until about 2015. Then someone with an actual understanding of the problem worked on it for a week and produced something that I would describe as poetry.
~Dijkstra (1988) "On the cruelty of really teaching computing science (EWD1036).
As a manager, I would value a developer that spent a week, refining a small, high-quality, robust and performant class, than one that churned out rococo monsters in a short period of time.
I tend to write a lot of code, and one of the things that I do, when I refactor, is look for chunks I can consolidate or remove.
OO is a good way to do that. It's a shame it's so "out of fashion," these days. The ability to reduce ten classes into ten little declarations of five lines each, because I was able to factor out the 300 lines of common functionality, is a nice feeling.
An interesting metric for me, is when I run cloc on my codebase. I tend to have about a 50/50 LoC (Lines of Code) vs. LoC (Lines of Comments).
Companies only want to grow, they don't care anymore about polished products.
Of course, I'm generalizing, there are some few companies that care about the product, but they are just a few.
It's not just twice as good to write shorter code. Its something like 64X as good. By some ways of thinking.
One bit of code I received was somewhere in the range of 5-10k SLOC when handed to me. I reduced it to around 1k SLOC.
The original included numerous duplications, had a function that was itself on the order of 1-2k SLOC, was miserable to extend (which path do I need to follow to insert this new conditional? which paths will be impacted if I remove this conditionals?). Fortunately it wasn't shipped code, but it was useful for testing the embedded system.
My refactoring involved, first, re-coding one of the larger, outer if-else-if sequences as a simple switch/case. This quickly revealed which variables were common to a large number of the branches and could be brought to a higher level of scope. Then common series of statements were parameterized and extracted to functions (preambles and postambles, setup and teardown, if you will). Several hundred lines were replaced with something like one 20-line function which took a single parameter and one call to it with the parameter that was previously evaluated through some hairy if-else-if sequence.
Once completed, new test sequences could easily be coded up or were available because there was already a path that led to it, you just had to feed it a different value or series of values to trigger that path. Then we could automate the whole thing because we knew what the parameters were that clearly led to each branch.
In the end the code was 5-10x smaller, but immensely more valuable.
I think the real reason is that as I moved to refactoring (as part of that 600 LOC experience), my deletions per year went up but my deletions per story regressed toward the mean.
No one had come close to that record in a decade and I beat it and you can kiss my keyboard if you think I was lazy about it. Harumph, I say!
is packed with stories which will make you smile, cry, or more enlightened. Or all of the above at the same time.
I'm starving for more stories of this size from CS history. PARC, Bell Labs, wherever! I'm sure there's thousands of fun little stories out there.
My personal favorite is The soul of a new machine, an account of a team trying to create a new computer : https://en.wikipedia.org/wiki/The_Soul_of_a_New_Machine
Compiled with what? Can one comply with an absence of request?
3/10 for realism