Harmful Gotos, Premature Optimizations, and Programming Myths
videlalvaro.github.io
videlalvaro.github.io
Very hard to analyze what was going on, very hard to refactor, very hard to reason about this program.
I was able to break out some parts into functions, and turn some gotos into loops, but for the most part it resisted my efforts.
This is what Dijkstra was complaining about - programs that resisted proofs because they were hard or impossible to decompose. He was not pedantic about goto.
You shouldn't worry about a goto here and there when it makes sense.
I had to do something like that once. I printed out the program on fanfold paper, took over a conference room with a long table, and traced out the control flow with colored markers. I was able to fix the program. That was not fun.
I have never coded a goto in any language since the days of FORTRAN and assembler. If I need to bail out of the middle of something, I'll use a break or a return, even if this requires encapsulating a loop in a function.
Note that languages with nested functions facilitate programming without "goto". When you need to encapsulate something which may require a bailout, it's easier.
Rust does nested functions well. The compiler can tell whether a nested function references an outer scope and thus needs to be a closure. If the function doesn't, an ordinary function is generated.
https://github.com/kstenerud/KSCrash/blob/master/Source/KSCr...
If I ever had to add more resources or steps into the mix here, I'd be in a world of hurt pretty quick if I were using a complicated nest of breaks and if-else. One could forcefully encapsulate each step into a separate function, but then you start to trend towards a forest of tiny little trees that you're not even sure what to name anymore, and the overall structure of the function and what it's doing become muddled.
Python's "with" clause seems to be the best solution so far. The "with" clause calls __enter__ at the start and __exit__ at the end. The __exit__ gets called no matter how you leave the "with" clause. This includes exceptions and returns. Importantly, there are parameters to __exit__ which provide information about what's going on when an __exit__ is executed because of an exception. Exceptions in __exit__ can be handled and passed on, even for nested "with" clauses. Writing a proper __exit__ is a bit tricky, but that's typically done only for standard resources such as files, connections, and locks.
So nested error handling via exceptions can work well if the scoped open mechanism ("with") and the exception mechanism play well together. It took a decade to get that retrofitted into Python, but in the end, it worked well.
Rust's error types are sound, but result in rather bulky code. I've written on this before.
This isn't quite right; a function declared with `fn` is always a plain function (i.e. never a closure) whether or not it is inside other functions. It's an error to refer to variables in outer scopes.
"Functions" declared with `|...| ...` are always closures, but there's been a few discussions about allowing such functions that capture nothing to be treated as function pointers too (not implemented or formally proposed, yet, though).
Aren't there tools which can help with this?
It used to be you could jump to anywhere in many programs. That would be understandably annoying.
But I've never used a modern programming language that allowed you to jump out of the current function scope. In that case, there's only a small range of lines your goto can go to, hardly what the original essay was talking about.
Now if someone writes a C program with thousands of lines in main() and goto as the primary control construct, then yes, complain about it.
Having done a lot of work in Asm, in which the main control structure is essentially a goto, I think it's something you just get used to; and often, I find that some algorithm which I've designed on a flowchart requires one or more gotos to implement in an HLL and I can't eliminate them without introducing more complexity, so they stay. In particular, algorithms based on state machines and related techniques tend to be quite "gotoful".
It's a bit sad that flowcharting isn't taught more these days, since I think it can lead to some really elegant and efficient algorithms and is, after all, how the machine actually works.
Also worth noting that Knuth's TeX, which is widely considered to be an example of highly efficient and bug-free code, contains quite a few gotos.
Regarding gotos, I think a more general statement would be "excessive reliance on one type of control structure is to be considered harmful." Avoid dogmatic preoccupation with general statements about avoiding language features.
Sounds a lot like FORTRAN IV. From what I remember of that version, that was pretty much all you had.
He would have only done better by doing that if his goal was to merely educate people that goto can be good too.
Which is very far from being his goal.
I have a copy of the original paper but the paper lacks the historical perspective of the post. I've never been able to find that blog post again. Does anyone remember what I'm talking about?
http://www.leeds.ac.uk/educol/documents/00003759.htm
http://stats.stackexchange.com/a/1516
And
Nassim Nicholas Taleb
Standard Deviation The notion of standard deviation has confused hordes of scientists; it is time to retire it from common use and replace it with the more effective one of mean deviation.
https://web.archive.org/web/20140115171745/http://www.edge.o...
Ironically, from the code I've seen, I don't think I've ever seen anyone actually misuse goto, probably because the people who use it actually know what they're doing. That's just anecdotal on my part, but still.
I've lived through the 80's, when 8-bit Microsoft BASIC was a popular programming language. Misuse of GOTO's was quite common, albeit unavoidable.
There was no lexical scoping in BASIC [1] (and all variables were global but that's another topic) so one could literally GOTO any point in the program.
[1] I'm talking about the BASIC one found on most 80s home computers at the time.
Ironically, from the code I've seen,
I don't think I've
ever seen anyone actually misuse goto,
probably because the people who use it
actually know what they're doing.
That's just anecdotal on my part, but still.
This no longer is just anecdotal. An empirical study on this just got linked to on Slashdot the other day.
http://developers.slashdot.org/story/15/02/12/1744207/empiri...I've misused goto at least once in the brave new world(tm) of structured programming. That said, once the cleaner way to do things occurred to me, I refactored it away immediately.
Occasionally when I feel I need something goto-like in very small code I find it is more cleanly expressed as a `for(;;){}` loop with either break or continue as the jump points.
I'm not saying goto is never a good idea. I just want to see a single good example.
The thing I look back on about "goto" and "gosub" is that, within the confines of what those simple control structures offered, much of my code wasn't actually structured that differently from what I would go on to write in more procedural languages down the road.
If you factored out most of my programs into pseudocode (allowing for limitations like global-only scope); the result would pretty much look like any QB or C program seeking to tackle the same problems, with main loops, clear procedural calls, etc. It was even idiomatic in those days to reserve a var or even a point in memory to take that gosub's "arguments".
The result of which is that when I did move up to QBasic, it was a pretty simple, easy transition, because it basically just allowed for a cleaner way of writing in the style I was already hacking together with DECB's primitive tools.
Good programming patterns emerge because they're the most efficient way of doing things. Goto is just another tool to do what a body needs to do.
Instead, ITT, gotos are bad and here's why, gotos are ok and here's why, gotos are good and here's why.
So sure, many of these things are mythical in scope and lore. But... the gist of many of them are not so misunderstood as to be wrong.
And then there's of course the greater narrative about the state of the industry, people chasing easy mantras instead of computer science, programmers ignorant of programming history, cargo cult programming etc.
in short, it's not a golden rule that some people seem to think it is.
I've never seen any actual mathematical proof of CAP - only a great deal of ( mostly well founded ) engineering conjecture.
Has anyone actually produced any mathematical proof of CAP?
Great question!
This presentation, linked by the article, explains the misunderstanding: https://www.youtube.com/watch?v=Wp08EmQtP44#t=1273
The paper you are looking for is this: http://dl.acm.org/citation.cfm?id=564601
Also available here: http://webpages.cs.luc.edu/~pld/353/gilbert_lynch_brewer_pro...
Many believe the problems are solved when nesting is concealed using named functions. It alleviates deep indenting, but it may make the code jump around while still dealing with shared state, propagating errors, and threading issues.
I have argued that callbacks are becoming the new GOTO statement-- useful in some contexts, but frequently overused in ways that were never intended. The result is "spaghetti code" that is difficult to follow and maintain. This paradigm has led to the popularization of the term "callback hell."
When developers totally ignore optimal code practices, bad things can happen. Things like 25000 (!) string allocations for each character typed in the Chrome "omnibox". [1] That's not likely one small bit of code with a few std::string's being allocated, but probably dozens or hundreds of (albeit minor) fixes spread throughout the code.
It is absolutely worthwhile to code optimally as a habit, when the optimal coding practice isn't significantly harder to implement or understand than the suboptimal practice. (In the above case, having a "pass all strings as const references" would have likely prevented most of the allocations.)
Similarly, the way you architect a product can strongly influence how easily it can be optimized. I've cut processing time down from 10's of minutes to under a second on some tasks, just by improving how the code was written -- and that approach means that when I write servers, I can handle all the likely traffic I'll ever see on a handful of servers instead of having to create additional layers of complexity to handle dozens of servers working in parallel.
So I disagree that "premature" optimization is a bad thing, by the typical definition of "premature."
[1] https://groups.google.com/a/chromium.org/forum/#!msg/chromiu...
In real terms, writing perfomant code is good, but packing data structures in interesting ways and optimizing in other manners when you aren't even sure you know all the data you need to store yet is often counter-productive.
Should you go down a rabbit hole, spending hours optimizing something because you think it might help? No, of course not. To my reading, that is what the warning is about.
But when you're designing a system, especially one that you know needs to be fast, should your design be optimized from the ground up? Yes, in some cases at least. [1]
When I code, I choose to use languages (C++, Go, LuaJIT, JavaScript/Node) that allow me to write code that ends up faster. Some people will cite "premature optimization" based solely on that decision, and they'll write their code in Python, or PHP, or Perl, or Ruby. [2] Since my servers will handle 3000 user interactions per second with no caching, I still contend that this is an appropriate stage for optimization.
I've seen projects fail because the server requirements were too high: 100 concurrent users per Python-based server on a free-to-play game meant that if the app took off, it would be losing money because of the costs of running hundreds of servers. You can't just "optimize" the app if the problem is you need to rewrite it in a completely different language. Computer time in the cloud may seem cheaper than programmer time, but in practice if you pay the programmers up front to do it right, you can cut ongoing operating costs indefinitely. And sometimes that can mean the difference between a project that succeeds and one that fails.
[1] http://gamesfromwithin.com/data-oriented-design
[2] It looks like at least Ruby has a server that uses a similar strategy to Node or Nginx+LuaJIT: http://www.akitaonrails.com/2014/10/19/the-new-kid-on-the-bl... -- but Ruby and Rails are ugly for a million other reasons, in addition to being slow, so I still would shun it.
That's true enough, though it doesn't seem to be the original point, since the context refers to optimizing parts of the code that aren't taking most of the time.
Which is a danger, but I think the danger is in wasting time on such optimizations, either up front or through increased complexity.
The reason I love it is because it offers really, really powerful ways of handling crazy amounts of abstraction. I'm absolutely willing to pay a performance penalty if it means I can, when I really start to need the performance, refactor the code in a day to where I can replace the needed bits with a C extension, and then spend the next day writing said extension.
Assuming a C extension is even called for. For something like a server, what I would do is, with Ruby, nail down the one thing that server needs to do, say, take requests and write them to a database. Then write that exact server in short and sweet Go or C++.
In fact, I can easily envision a time in the not-too-distant future where the broad strokes of my coding practices are set down and I'm down to refining them so they work faster and better, working towards a clean mix of Ruby and C where I can refactor concepts into and out of both languages as needed.
With Ruby, I can spend less time implementing and more time thinking hard about my domain model. Ruby is not ugly to me, it's incredibly beautiful, particularly its object model. It's great because I can write ugly code to make something work with an eye towards making it easy to clean up later. When I write something ugly, I know that there's a domain concept that is going to need to be let out at some point, the wonderful thing is that I don't have to do it immediately. Ruby lets me manage a monstrous mess of interrelated concepts and abstractions, manage the process of starting ugly and iterating towards clean. And I can do it all by myself.
I see nothing wrong with choosing to start with a faster language if that's where your skills lie. But you can pry Ruby from my cold, dead hands. Does Ruby scale? I wouldn't want to try with anything else. It's not just connections per second that needs to scale, it's everything else too.
But I write code in LuaJIT that's typically as clean as or cleaner than the equivalent Ruby code, as well as already approaching the speed of C, so I don't even need to rewrite it to be performant later.
But by all means, if the tool works for you, use it.
IMO these days even the first is bad advice if taken literally. You need to ensure you don't work yourself into a design that cannot be optimized without being rewritten, which is unfortunately a problem I've seen a lot.
You need to do design-/architectural level optimizations from the start.
Of course, the "should we use goto" question is different from "should we implement goto in new programming languages".
However, few people actually do: https://peerj.com/preprints/826v1
there are no good/bad features, there are only good/bad programmers.
If we follow this to its extreme, everyone should switch to Brainfuck.
The complexity contributed by a feature is partly a function of how unintuitive it is. "Goto" is so intuitive that it's essentially free, in the sense that it doesn't really increase the cognitive burden of the language. (If misused, it can greatly increase the cognitive burden of individual programs in the language, of course.)
Yeah, we shouldn't.
Also, one of the things I found actually reading TAoCP is that in pretty much all algorithms, he gives a full instruction count of everything. So, it isn't like these preclude each other.