I hate cut-and-paste
jacquesmattheij.com
jacquesmattheij.com
I worked on a large embedded system where all the principal developers used such an editor (there was some occam heritage). One incautious copy and paste of 'a handful' of lines, and you had accidentally replicated a complete subsystem! It made debugging 'interesting'.
Every time I think 'Oooh - code folding - what a great idea', I retrieve those memories.
More generally, any time editor cleverness is interposed between the creative idea and the source code, I get a twitch. It should be possible to work on the source code without requiring anything more than a bare-bones editor - that is all I can assume about what my favourite editor shares with whatever future developers (incl. myself) happen to be using.
But to complete Jacques' rant, I'd say that if duplication is bad in the part of the code handling logic and actions, it is not bad and often "normal", or even "better" in configuration parts or in content parts (eg templates). Factorizing a list of configs to replace some duplication with logic (loops, conds) can be a very bad idea.
Problem is copy-and-paste, which litters codebase with duplication and lead to messy and hard to maintain code.
> Think about that ctrl-C, ctrl-V sequence next time you are about to use it.
Many people use "cut and paste" as a generic term that refers to either copy and paste or cut and paste. You have to figure out which they really mean from the context.
I currently hate an application that I'm likely to end up working on that was authored with a generous helping of copy-paste.
What's great about ed is that you keep most of the program in your mind. There is no tab completion. There is your brain. And when you keep the whole program in your brain, you write less code and the code you write is better thought out. I wish we could go back to the days of ed, but all we can do now is consciously resist the desire to work at our tools' maximum capacity.
I have been wondering that, actually.
On the topic of the article, luckily, the Rails community (and python, and others) seem to have embraced this idea, under the heading "DRY principle" (Don't Repeat Yourself). There might be a connection between that embrace of DRY and the fact that Ruby/Python afficionados tend to prefer "simpler" text editors (Textmate, Vim, Emacs, Sublime) rather than IDEs.
Here's an example from when I was very first starting out with ruby:
This is a pretty hard to read piece of code, despite my rambling comments. All done to avoid copying and pasting the same easy-to-follow method ~10 times.
(Could still be me being a noob, but I can't see some other architectural means of avoiding a choice between either duplication or tricky-to-read metaprogramming in that case.)
It's shameful and I'm sick of it. I want out. Right now, I'm paid quite well as a contractor to make some enhancements to a project developed by a hundred underpaid developers in a distant land, who did a remarkable job, considering. But it's rat's nest -- they didn't stand a chance. And I can't look myself in the mirror anymore and say this is what I do. It's grotesque. I'd rather run a coffee shop (and indeed I've been talking about it) because at least I could go home at the end of the day having seen a little beauty here and there, instead spending all day staring into the face of the Elephant Man.
1) Sure it's ugly, but look at what it can do!
2) Ok, it's ugly, but with good practices, good people, good management, you can ameliorate the worst of it.
3) Fuck it.
I've decided that this is my last contract. I have a project called Kayia, and my wife has a site called Kongoroo. I'll continue to work on those but otherwise that's it. If I'm not coding in Kayia, I'm not coding.
It doesn't sound like burnout to me either. I went through something similar. The solution was to admit that I had taken a wrong turn in my programming career, and commit to working only on things I believe in. It was either that or get out of the software business altogether.
http://c2.com/cgi/wiki?BurnOut http://c2.com/cgi/wiki?GetaLife http://c2.com/cgi/wiki?JustLeave
I recently stumbled on them. Especially the second one has some entries that are very inspiring if you are feeling burned out / largely incontent with programming.
It's not just where I'm working, etc. -- it's all based on C. It's all a farce. (And don't mention Lisp, which is an idiot savant. Sure it can count the matches on the floor, but you can't introduce it to anyone.)
(p.s. Edited for brevity. I always do that and never bother saying it, but in this case there was a concurrency conflict because someone downvoted the bloated version. Mea culpa.)
I remember a lunch I had with Kent Beck in Zurich in 1998 (I think) and I was all set to pounce on him with this idea I had of a new approach and he instead pounced on me first with XP. I was utterly unconvinced, not because he didn't make a fantastically convincing argument, but I was -- and am -- convinced the failure stems from somewhere else.
Another way of putting this critique is that XP is designed to work inside existing companies that are incapable of tolerating the forms of organization needed to produce software well, when what we ought to be doing is starting new and much smaller organizations, which in the local dialect I believe is called "startups". Of course it's not that easy, because once startups begin to grow, they bring back the old organizational assumptions. But that just means we need more deviant startups.
You, on the other hand, think there exist paradigms of abstraction that can fit the complex systems we're trying to build well enough to make the process tractable. I hope you're right. Certainly some such forms are better than others, so others could be better still.
How do you propose to demonstrate it?
You've summarized my argument nicely.
How to demonstrate it? Are you asking me to put my money where my mouth is? :-) Good question. I'm working on it. For example, we need to query code. I want to see all aspects of this portion of the UI. We should only look at code in layers, as accounting systems do. For one. I could go on, but I'll leave that for sometime over a beer.
This one, you may be able to introduce to people (of course, OMeta by itself isn't enough: the DSLs you could write with it however…).
Often it's a good idea to start by cutting and pasting, and then afterwards figure out what ended up being sharable.
Seriously though, this article got me to thinking and I realized that why would one live with the smell of copy-and-paste when it's just so darn easy to write a singly reusable function?
I agree with what you in the sense that there's no need to go from what might be a few lines of copied code to a full-blown library or subsystem. But for a small bit of common code not refactoring that out immediately seems to me a really bad practice.
The huge abstraction with class hierarchies is an entirely different kind of idiocy. It's orthogonal to this problem. You can (and I've seen it done) cut and paste giant class hierarchies too.
(Edited) Also: http://clonedigger.sourceforge.net/ Use it for great good.
I would expect such a tool to parse the language into AST form and find branches that are the same except some identifiers and a few other details. It is probably intractable in general, but I think it is feasible for most code bases.
Huh?
Turned out it was a copy/paste job... which in turn became only the start of the rabbit hole. :)
Most of it is throw-away code: just playing with ideas and different implementations. Exploratory programming is cheap these days, and I really think it's for the better.
If it looks like I'm narrowing down on something I'm actually gonna use, I refactor, rewrite, simplify, and delete a lot of code. ctrl-x becomes my new best friend. (I write in natural languages much the same way.)
So many times, when you need a quick solution for a small problem and you know that you have fixed this exact same problem a few weeks/months ago at a different spot, there's a huge temptation to just go and copy&paste those lines.
By doing so, you have just created debt. What if that code is using some API that you want to change a few years later? Whoever is going through the old code now has to fix up all instances of your copy & paste action.
What if the initial code contained a bug that somebody else fixed? They might not know about the copy pasta. It's very likely that they only fix the initial instance of the code and not all other places where it was pasted to, so the bug partially remains, or, worse, is later classified as a regression (which it's not).
So, coming back to your analogy, hating copy & paste is analogous for hating guns not only for their potential of killing people but also for their potential for causing accidents and for their potential to use them for any kind of potentially non-violent crime.
As the guy who is usually doing the refactorings to our 7 year old codebase, I'm the guy who is suffering from copy pasta and I'm telling you: I totally agree with the original article. I've yet to see a single instance where copy & paste of more than a single line of code didn't cause me non-insignificant amounts of additional work.
To get code shipped I usually copy and paste if it really saves me time, but make a checklist of optimizations I can do at a later date.
When that time comes, find/replace does the dirty work.
What he really hates is his lack of willpower when it comes to doing things the wrong way. Cut-and-paste enabled him to be lazy, but it certainly didn't force him to be.
But like any sharp tool it has the ability to cut you just as easily as it cuts the wood, so you need to apply it with care.
Like all tools, it has its place, and there are people who tend to misuse it. Just because it's misused by bad programmers is no reason to deprive GOOD programmers of its benefits. You might as well ban screwdriver heads in power drills because lots of clumsy, inattentive people tend to strip screw heads with them. Similar misguided calls-to-arms have been raised over other useful tools such as preprocessor macros and goto.
Okay, there are some (corner) cases where they really are a good idea. But "never ever use this Chtulu Abomination" still is a damn good heuristic.
Preprocessor macros: Code generators, compile-time switchable code (such as logging) without filling your source files with #ifdefs
Goto: Managing complex resource allocation/deallocation within a function:
bool success = false;
int fd = -1;
char* memory = NULL;
fd = open(filename, "rb");
if(fd == -1)
{
LOG_ERROR("Could not open %s: %s", filename, strerror(errno));
goto done;
}
memory = malloc(BUFF_SIZE);
if(memory == NULL)
{
LOG_ERROR("Out of memory");
goto done;
}
for(blah blah blah)
{
if(failed to read or whatever)
{
LOG_ERROR("It's borked");
goto done;
}
}
...
success = true;
done:
if(memory != NULL)
{
free(memory);
}
if(fd != -1)
{
close(fd);
}
return success;
I disagree that dismissing a tool out of hand as an "abomination" is a commendable approach. If you take that attitude (or instill it in others), you'll probably never learn the proper use of such a specialized tool, leaving you ill-equipped to deal with the situations those tools handle well.Unrolling loops: I never saw an instance where that was necessary. Plus, the compiler can often do it for you. I know we often use high performance applications (video decoders, 3D games…), but very, very few of us write ones.
Switch statement: the syntax of the construct is heavy, we should lighten it. The rest is hardly boilerplate any more:
light_switch (expr) {
1 : i11; i12; // "break;" is implicit
42 : i21;
: i_default; // "default" is implicit
}
Preprocessor macros: I agree (I mean, I back-pedal), they are more useful than the rest. However, I still avoid them by default, as they make really good foot-guns.Goto: your example shows exceptions (try…catch finally here). Goto makes much less sense when you have them.
Now my point isn't to never do those things at all. Only to think of them as last resorts. The "Chtulu Abomination" metaphor helps me do that.
All languages have boilerplate somewhere. It's unavoidable. Switching languages just because there are pain points is not a solution, because you'll simply be exchanging one problem for another.
"or modify the front-end of your compiler"
This is most definitely pie-in-the-sky. In the real world of real business, you can't do this.
"I mean it, the cost of boilerplate is really high."
Oh, I agree wholeheartedly. But you need to work in the languages that programmers understand today. I'm not going to have nearly as much success hiring smalltalk programmers as I would have hiring Python programmers, for example.
"Unrolling loops: I never saw an instance where that was necessary."
I have, but then again I've been at this for 20 years.
"Plus, the compiler can often do it for you."
You can't be sure until you look at the disassembly. Often, it does it wrong.
"Switch statement: the syntax of the construct is heavy, we should lighten it."
That's not going to happen within the next decade. Meanwhile cut & paste is my tool of choice to cut through the boilerplate.
"Goto: your example shows exceptions (try…catch finally here). Goto makes much less sense when you have them."
But it makes LOTS of sense when you don't have them. And it's still useful in C++ and Objective-C, where you still end up interfacing with C libraries. If you don't know about the goto "poor man's exception", you'll either end up with repeated and buggy deallocation code, or a monstrosity of nested scopes, or a bunch of check-return-and-throw constructs, which in the case of Objective-C will slow your program way down because it implements try/catch using longjmp.
Goto is also useful for breaking out of inner loops when the language doesn't support that (some languages support "break [label], to make it seem less like a goto).
I don't consider these to be last resorts; I consider them specialized tools. Much like design patterns, they're for mitigating some deficiencies in a programming language, but they're only effective if you know how to use them.
However, I'd like to challenge the belief that modifying the front-end of a compiler is too hard, or unreasonable. Even in the so-called "real world" where screwing up means you're fired.
First, the point is to tweak the language, not the compiler. For instance, we may want to lighten the switch() syntax before GCC does, but we do not want to modify GCC itself if there's a simpler way.
More often than not, there is as simpler way: just write a parser and a printer for your language, so you can do source-to-source compilation by chaining them. Printers are easy. Parsers are almost as easy, except for C++.
Then tweak your parser (lighten some syntax, add some keywords…), do some pre-processing between the parser and the printer (yeah, true macros), whatever.
Now there are some caveats: such a pre-processor may confuse IDEs, and may screw up error reporting (where errors don't track back to the actual source code). I personally don't care much about the former, but to solve the latter, the base compiler need to provide a way to be told where a given line of "source" code actually comes from (very useful for tools such as Lex/Yacc). Unfortunately, taking advantage of this will greatly complicate your pre-processor.
Thanks to eclipse I go days without copy-pasting. May I suggest that if you're copy pasting too much you're doing it wrong?
If we take it to the extreme, why would you need to have the same code in two different places?
We tend to follow the path of least friction. If a tool you use changes that path, and unless (even if?) you consciously fight it, you will change your behaviour, whether this is a good thing or not.
Please, please learn to write correctly. "IDEs", not "IDE's"