-2000 Lines of Code (2004)
folklore.org
folklore.org
At a previous job, a pretty smart architect decided that Sonar was a pretty good tool to gauge code quality. He wasn't wrong. However, as a way to improve code quality, management decided to make the continuous integration server demand 90% test code coverage minimum, including branches, and no loss of code coverage over half a percent from maximum, as measured by Sonar.
So I inherited a large, rather terrible application, which included a whole lot of UI code using Swing. Swing has it's qualities, but being easy to write meaningful tests for isn't one of them. So my predecessors often left some chunks of frontend relatively untested, while making sure large amounts of business logic was well tested. So to keep the application in order, I started doing a bunch of refactoring. I cut the size of the backend in half, and all the tests kept passing. But my commits were rejected by the build.
So if you have 10K lines of untested code, and 90K lines of very ugly, badly factored, but tested code, you just couldn't remove 10K of the bad, well tested code, because that would sink coverage metrics.
After much arguing for relaxing the rules, as I didn't think that adding tests after the fact to bad UI code seemed like a good use of our time, I fixed it in the only sensible way I could: Added an extra ten thousand lines of well tested code that didn't actually run in production, but made the metrics happy. Only then I could get the build to pass again.
"Measuring software productivity by lines of code is like measuring progress on an airplane by how much it weighs."
There seems to be quite a lot of lines of code there.
http://ianmurdock.com/platforms/on-the-importance-of-backwar...
Original ASP.NET blog entry was gone, best I could find.
Oh yeah, that's Moore's Law at work here too!
You're throwing away the top end of the funnel so it's impossible to know whether it's leaky or just completely dry all the way through.
The solution to either problem in those examples is the same. More engaging content will drive more visitors to the site as well as stimulate more conversation.
However I do agree with you in theory about how it's good to keep a log of traffic; with the caveat that it's not used as the primary measure of activity when conversation is end goal of the site.
My point today is that, if we wish to count lines of code,
we should not regard them as "lines produced" but as
"lines spent": the current conventional wisdom is so
foolish as to book that count on the wrong side of the
ledger.
[0] http://www.cs.utexas.edu/users/EWD/transcriptions/EWD10xx/EW...With the sheer amount of bad code people write, I expect to do a lot of deleting, refactoring, and rewriting, and I'd hope managers/fellow team members would be able to see the value in that. But sadly, they usually don't.
git log --author=rav --numstat --no-merges --pretty=format: |
awk -e '{ a += $1; b += $2; } END { print a; print b; print a-b; }'
produces this output (sum additions, sum deletions, net line additions): 112529
85383
27146 git log --numstat --no-merges --pretty=format:%an
| awk '
author == "" { author = $0; next }
/^$/ { author = ""; next }
{ added[author] += $1; removed[author] += $2 }
END { for (author in added) {
print author, "added", added[author], "removed", removed[author], "sum", added[author]-removed[author]
} }'
| sort -n -k 7It wasn't really all that tricky though, it took me a few hours to write. git-log has options for only displaying the status line of diff-stat for each commit, and then displaying the parents of each commit, and the author. You look to see that there's only one parent (so it's not a merge), parse out the X added, Y deleted numbers, and stick them in a dictionary keyed by name.
A lot of the script was just getting statistics like average/min/max/stddev line counts, and printing them nicely.
I'd say it is open season on commented-out code. If code has been commented-out since a few weeks it is good to delete it just as a matter of course.
Of course this leads to nasty conflict resolutions, so they solve that by squashing all of their commits before rebasing.
Now your code is no longer in the repo.
Hell, we keep our feature branches around here. *shrugs
And I've done similar things with "undocumented compiler directives" for XSLT processors before. If you didn't leave the comments in then you got the dreaded "GregorSamsaException". I refer to it as "dreaded" because folks on our team dreaded it... the exception occurred at unpredictable times and who the hell was Gregor Samsa anyway, and what did that have to do with our Java application?
Turns out, Gregor is the main character of Kafka's "The Metamorphosis" and the exception was the brilliant idea of someone who wrote the XSLT transformer and probably thought it was cute. (It wasn't.) It was an internal exception in the transformer. (Get it? Metamorphosis? Transformer? Never mind.) It occurred when a certain buffer filled up, and the exact length of the XSLT input file affected that, so adding a few lines of commented-out content would make the error appear or disappear.
I wrote a lengthy essay explaining the above facts, and used commented-out excerpts of that essay as padding in the file. Yes, I was trying to be "cute" also, but I was too young to realize that was a bad idea.
I can imagine there being instances where leaving commented out code could be helpful -- including comments about why you thought it could be helpful, and why it is commented out!
That said, I do see where you're coming from. Part of the issue is that we (well, most of us) don't have good ways of searching old code. There is Codeq ( https://github.com/Datomic/codeq ), which is prettydamncool™ ...hopefully we'll start to see more systems like it.
I've increasingly been noticing that a lot of really good development practices make sense if you're starting from scratch and can employ them right away, but sometimes if you've got years or decades of legacy code and legacy process (and code that was written as the result of legacy process) to deal with, the right thing to do isn't always so clear.
(I've unfortunately/fortunately been working some recently on a very large, very old code base that mostly doesn't need updating. Trying to unravel its mysteries enough to add a new feature has been an interesting experience.)
But how will you know that you should look for it in the first place?
LibreOffice put the code into git and went mad with an axe deleting all the commented-out code.
Apache OpenOffice, on the other hand, still commits new commented-out code. http://mail-archives.apache.org/mod_mbox/openoffice-commits/... Possibly nostalgia for the good old days at StarDivision.
Also, even more occasionally, I'll leave some incomplete code commented out as an obnoxious reminder to complete it later. This is especially useful if the code wasn't ever committed before, so checking it out via source-control isn't an option. The very fact that it's not really supposed to be there is a good motivator to implement it!
I'll admit to just commenting and uncommenting log lines before, though - learning how the logging systems of major programming languages work takes some time, and there's an up-front cost to starting to use them.
It all really depends on the application and how much performance is an issue.
I have little embedded systems experience, but what I gather from talking to folks who do is that they also use logging APIs, but they avoid logging from inside a hot inner loop (their log statements are usually around startup/initialization and when the system receives certain inputs), and they use custom log writers that write the logs to a host server or desktop system when plugged in for debugging rather than taking up storage space on the device itself.
Debugging is a highly interactive art and sometimes operates inside a much faster feedback loop than that. Put those few log calls in the innermost loop, run it, scan the 20MB dump for a weird entry, fix, remove logger calls, and you're done in less time than it takes to consider which key invariants and exit conditions are worth tracking.
For now, I can post a link to my inspiration: http://wordaligned.org/articles/cpp-streambufs
Extending from there is pretty straightforward, albeit you can hit some dark corners of C++ (I spent a few weeks tracking down a double link error caused by not templatizing an addition to the Logger that gave the capability to output to MSVS's debug window). This is also one of those very few cases in which I have justified using multiple inheritance, virtual inheritance, private inheritance, and templates.
I don't think that a commit should be littered with commented out code, but there clearly are positive reasons to have some.
Besides, if the comment is not there, how can you be expected to know that there is old but still relevant code in the repo?
# TODO: Figure out why do_foo is triggering a bug
# http://my-company.org/issues/42
# def do_foo(self):
# ...If the code is in past commits, there isn't a reason to muddy the source tree with it.
[1]: http://www.folklore.org/StoryView.py?project=Macintosh&story...
No internal slots, isn't this Steve Jobs' dream machine?
Some said: "A piece of art is only finished, when there is nothing left, that can be taken away."
The art in computer programming is, to find ways to bring the problem to the point, to find out what is really necessary to solve the problem (and not more). This reduces (often, not always) the runtime, the amount of memory needed -- and (most importantly!) the amount of maintenance that is needed. The maintenance of a program is directly dependent on the number of code lines.
Many companies start fast with a superior product, but than comes the time of growth and growing demand, new employees are rushing in ... and the number of lines explode. That is the point of danger. The company is about to strangle itself. The number of errors are rising.
I remember an old, but once famous database product. The first 2 or 3 versions where great and the company grew out of 4 developers to a horde. The next version came out much much later than expected and was first a bug ridden chaos. The problem: The number of employees and the number of code lines grew faster than the company could manage them.
It is also said: "Adding new members to a late project, makes it later" That's because of the overhead of managing those peoples and the added code does not always add to project speed.
Then again, "A work of art is never finished, it is abandoned." (http://www.quoteyard.com/art-is-never-finished-only-abandone... )
The database product sounds like dBase. If so, certainly management focus was also a factor. http://en.wikipedia.org/wiki/Ashton-Tate#dBASE_IV:_Decline_a... attributes it to a push for Diamond. Its source material is unattributed. While http://www.dbase.com/Knowledgebase/dbulletin/bu03_b.htm says it was a management push towards OS/2.
Your last quote is from "The Mythical Man-Month" by Fred Brooks(1975) and is called "Brooks's Law".
More seriously, of these two implementation for Python's str.lower(), which is "minimal"?:
for (i = 0; i < n; i++) {
int c = Py_CHARMASK(s[i]);
if (isupper(c))
s[i] = _tolower(c);
}
for (i = 0; i < n; i++) {
int c = Py_CHARMASK(s[i]);
s[i] = _tolower(c);
}
I suspect most would say the second is minimal, but the first is what Python uses, because it's measurably and distinctly faster than the second, and that extra performance is worth the extra maintenance overhead.That's one reason I believe that programming-as-art metaphors don't really apply. Minimalism as an art form not the same as minimizing the cost function of different uncertain factors.
Anyway, this reminds me of a story from the Tao of Programming, http://www.canonical.org/~kragen/tao-of-programming.html :
There was once a programmer who was attached to the court of the warlord of Wu. The warlord asked the programmer: ``Which is easier to design: an accounting package or an operating system?''
``An operating system,'' replied the programmer.
The warlord uttered an exclamation of disbelief. ``Surely an accounting package is trivial next to the complexity of an operating system,'' he said.
``Not so,'' said the programmer, ``when designing an accounting package, the programmer operates as a mediator between people having different ideas: how it must operate, how its reports must appear, and how it must conform to the tax laws. By contrast, an operating system is not limited by outside appearances. When designing an operating system, the programmer seeks the simplest harmony between machine and ideas. This is why an operating system is easier to design.''
The warlord of Wu nodded and smiled. ``That is all good and well, but which is easier to debug?''
The programmer made no reply.
Rephrasing for the software world: "...without having to make appreciable sacrifices in user experience."
Recently I've been learning that from two angles. First I'm realizing that being a bit more verbose is often clear and safer. Second, I've been reading and learning a lot about compilers, and I'm realizing that my "efficient" (short) code and other trickery does diddly-squat, and the compiler will produce the same code either way.
I once sped up an intern's code by just deleting a 30-line function, and doing nothing else. Doofus didn't realize our language had a highly optimized built-in sort, and so he wrote his own inefficient sort (insertion) that overrode the existing one. Poor little guy was so proud of having chosen exactly the right sort, and then implementing it based on his recollection of college... I told him to spend the next few hours just reading the documentation of the core API for our language.
Unfortunately, I forget who said it.
Isn't that backwards?
I can say this having spent months in the laboratory to learn things which were already known. Who knew that precision quartz pressure gauges were called 'manometers' in the late eighteen-hundreds? I didn't, and thought I'd invented something new for a couple months in my first year of grad school.
The adage is a good one. Reading a lot and reading widely pays off.
Who's to say an anecdotal 'employee' is any better than his 'intern'?
I was riding in a car with buddies who were much better C programmers than myself and making notes on some code I needed to do the review of the next morning. This is the code pattern that almost crashed our car:
a = some_function(i);
a1 = &a;
a2 = &a1;
calc(a2);
with calc() reversing it outI was cussing a bit much[2] and front seat passenger had to look then driver got too curious. Sadly, this was the least "wrong" thing about the code. Its very hard to do a code review where you suggest 500 lines of C can be reduced to 50.
1) cannot have the guy doing code reviews actually coding, that would be improper
2) cussing in private allows positive, motivational tone in public
― Antoine de Saint-Exupéry, Airman's Odyssey
I'm longing for the day when people realize how significant this finding potentially is.
Like any other powerful ignored truth—if that's what this is—its path to acceptance will likely be through somebody doing something impressive with it that hasn't been done before.
"Hey, I know you don't like this paperwork. I don't either, but I get asked to do it all the time. My strategy? Give them as much paperwork as I can. These forms, those memos, CCs on emails... eventually, sometimes, they ask me to stop. 'You're giving us too much paperwork', they say. 'You don't want me to give you paperwork? Works for me.'. So fill these out, and I'll take care of the rest. Now, for today's class..."
"One of my most productive days was throwing away 1000 lines of code." - Ken Thompson
exceptions are things like game programming, where certain hairy-looking tricks can be necessary and have a real benefit. can't say the same for your average web app.
Anyone who has completed a complicated functional program would probably understand how that could be a misleading measure of progress and, worse, lead people to spend time writing poorly engineered untested code. I think the people who implemented this system understood that counting lines of code was a dumb idea, but somehow thought this was better.
Congratulations Microsoft. You've now become the bloated mid-80's IBM you used to hate.
I think programmers nowadays are a bit overly obsessed on getting lines of code down
Its certainly possible to write something in one line when you could use 6 if you use the more obscure and "fancy" tools available in the language you are using. One thing that randomly springs to mind is linq in .net Reasons why more lines might be better.
1. The longer more explicit code might end up shorter and more performant after it is compiled (like in the oldskool performance increasing trick of unrolling loops)
2.When you come back to your fancy code later you may have forgotten about that particular fancy trick and now you dont understand your code.
3. Other people are less likely to understand your code.
4. By using more specialised features in a langauge your code is now less transportable to other langauges.
Personally I think the obsession with fewer lines quickly becomes counter-productive and the main reason it is done is in order to show-off your knowledge of these fancy things.
Note: I'm not talking about the guy in the article btw. Just talking a about a general trend I've noticed in modern programming.
From what I've seen, bloated code tends to be code that is copied (either literally or in style) from somewhere else, which doesn't fit the task at hand. So you have functions that have lots of options, but only ever get called once with one set of parameters. The fix is to rewrite everything so that the abstractions are taylored to the task at hand.
What you are describing is a kind of excessive elegance. But this is fairly independent of the problem I described (and the article is describing).
I'll also say that there is a limit to bloat, and LOC is often a good guide to functionality. E.g. if I'm interesting in re-implementing something, I'll first look at the weight of the code. If I thought something would be a half hour job and the code is 5000 lines, I'll probably re-evaluate. Sometimes I've pondered reimplementing certain libraries, only to find they weigh in at a million lines of code.
Or it might be less performant. If you actually care about this, you should have an automated test running regularly that will tell you one way or the other. But most of the time it's not worth worrying about such things.
> 2.When you come back to your fancy code later you may have forgotten about that particular fancy trick and now you dont understand your code. > 3. Other people are less likely to understand your code.
Or using fancy tricks more frequently can help you remember them. I think you should use every feature available in the language - developers need to be able to understand the language so that they can read third-party library code. Or else have an automated system that flags usage of particular features.
> 4. By using more specialised features in a langauge your code is now less transportable to other langauges.
Who cares? Seriously, how likely is this to actually come up? If you've chosen language X you presumably had a good reason for doing so; you should write language X, not try and write language Y in language X.
Fewest lines of code is not the perfect metric, but I think it hits the sweet spot: it's very simple to calculate, and captures a good proportion of the difference between good code and bad code.
A rant probably isn't one of them.
Were I to work in a place with a 'git commits' metric for productivity, I'd happily commit each line individually, to ensure that any potential data loss was as limited as possible, of course.
To paraphrase Wally from Dilbert: "I'm writing myself a mini-van."
>if you come up with such a metric and have a developer who can't figure out a way to game it, perhaps you want to consider letting that developer go.
Maybe that would be a good interview question (for a software manager): "How would you game metric X"?
Bah. While where I work has never even flirted with that metric, general agreement in my team was that the first thing to be done in the event that ever changed was write a git filter that committed one character at a time. I suppose you could take it further down to one bit at a time if you really wanted to; once you had the code for one character at a time, one bit at a time would be a trivial extension.
Less humorously, I've been complaining that our email notification system sends out a separate email for each commit in every new git branch... that is, if you are on a branch with 100 commits, and you take a new branch and commit that branch onto the server, our emailer seems to believe you just made 100 commits, and sends 100 emails. Or 1,000, as the case may be. One of these days that thing is going to take out the entire corporate email system.... of course, one never fixes the problem until it reaches that state, so I'm just waiting....
Making small atomic commits is actually good practice.
And since I was formally banned from using generics and lambdas and from refactoring old code without explicit permission, I am no longer hurting the team! Yay!
This is the worst job I have ever had.