In the Defense of Spaghetti Code (2013)
250bpm.com
250bpm.com
I think the only point of the article is "Don't try to retro engineer the logic behind a complex domain that you do not master".
It is fortunate that it is possible to show, via Turing completeness, that spaghetti code cannot compute anything that structured code cannot, for if not, we would still have programmers insisting (as some did, in response to 'Goto Considered Harmful') that spaghetti code is sometimes necessary.
> I think the only point of the article is "Don't try to retro engineer the logic behind a complex domain that you do not master".
AKA Chesterton's Fence:
https://en.wikipedia.org/wiki/Wikipedia:Chesterton%27s_fence
Also, I always think Turing Completeness is a red herring in conversations about Software Engineering. Turing Completeness is such a low bar, that it hardly factors into most conversations about the expression problem. It's most notable when the goal is to actively restrict completeness.
It's kinda like starting a conversation about birds and first establishing they have mass.
And I'm just pointing out that in this particular case, Turing completeness was very useful in decisively stopping what would have otherwise been an unending and inconclusive argument, and, regardless of whether or not it can justifiably be called a low bar, it was good enough in this case. Furthermore, It does the same in many other cases: for example, in the question of whether there are ISAs that are more computationally powerful than others. It seems to be, or at least have been, very useful in computer science even if it does not have much relevance to everyday coding, and it may not seem much of a big deal because the questions that it has solved are no longer (or never became) problems.
[0]: https://en.wikipedia.org/wiki/Turing_tarpit
> In any Turing complete language, it is possible to write any computer program, so in a very rigorous sense nearly all programming languages are equally capable.
You seem to have more history on the matter than I do, so let me ask a question instead trying to defend that idea. Why did anyone think Structured Programming wouldn't be? They must have had reasons to suspect it. I also don't know much of the history of Structured Programming apart from "Goto Considered Harmful", so pointers would be great.
You don't need to create function calls just to add comments.
int new_sequence_number (int old_sqn, int foosoft)
{
if (foosoft == 2) // version 5.34.2 of the protocol
return old_sqn + 3;
else if (foosoft)
return old_sqn + 2;
else
return old_sqn + 1;
}
into this: int new_sequence_number (int old_sqn, int step)
{
return old_sqn + 1 + step;
}
The problem is that this isn’t reasonable. Though the function signature looks the same, foosoft is an enumeration of some kind but step is actually an integer— This sort of change requires an audit of all call sites, as the meaning of one of the arguments has changed.One big red flag here is that the new function body isn’t what you’d expect from its signature. Why is the sequence number incremented by step+1 instead of step? The only reason I can think of is that they want to maintain compatibility with an old callsite, but that doesn’t make sense with the category change of the argument.
Ideally, the function signature itself should change so that the compiler will point out all of the obsolete usages to you. If you can’t do that with the type system, change the function name.
Because the problems now cut across responsibilities of the components built on certain assumptions made for the sake of removing the spaghetti. Because now you need to account for "global" parameters (protocol version) in lots of disparate "local" scopes - individual functions who are quickly losing their degree of encapsulation.
We're discussing a hypothetical code-base, so focusing on the details isn't probably going to be too helpful. In any case however, I in no way see how the argument "flows from the outcome" of the refactor you're focusing on.
Specifically, he’s using the failed outcome of this hypothetical refactoring to argue that new_sequence_number should never have been broken out into a separate function in the first place because it separated the foosoft==3 case from the conditional, which made the bad refactor more likely to happen.
Certainly, in this hypothetical 1500-line function, it’s possible that they would have appeared close to each other and it’s possible they wouldn’t have. In this particular case, it’s entirely possible that the refactor was made easier to execute successfully because the programmer can leverage the compiler to find all the locations that need to be updated.
In the end, he’s demonstrated a flawed hypothetical procedure and asserted that it’s worse than an imagined alternative. As usual, the argument being unsound doesn’t mean the conclusion is incorrect, only that it hasn’t been demonstrated here.
Personally (anecdotally), I don't believe that this statement is true. In my experience, an appropriate refactor will always be more readable than any spaghetti code. I may be wrong, your statement may be true, but without good examples of such refactors, who can say.
The above commenter is just pointing out that the example given in the article is absurd. Which it is. And with that, the argument is lost.
(And that one person proving me wrong and reading the 10,000 lines would then proceed to nitpick endlessly on yet more details irrelevant to the blog post's point....)
1500 lines of code for a function/method would strongly suggest a problem. In the same way as a 200 function/methods would make it hard to follow.
It saves you the job of inventing descriptive procedure names as well as jumping between procedures.
I seem to recall also that in Code Complete, Steve McConnell claimed that there was an inverse relation between the length of a procedure and the number of errors per line of code.
We have an informal, internal audit of our code base to highlight problem areas, trends, and note if something we're doing is problematic over time. Someone suggested a hard cap of 30 lines per function.
I immediately had a rebuttal with code. We have some functions that are mostly mappers. Ingest code from one place, transform/cast/sanitize/clean it, then send it elsewhere. The types of information for some of these functions can make some of these behemoths 200+ lines. I asked my peer if they would rather follow the trail back-and-forth of 20+ helper functions or look at one monolith knowing that it only maps data?
They wanted helper functions.
To counter, I highlighted a mistake that I made years ago. I traced the revision of a file where a function mapped data that used helper functions. I pointed out that I did the exact thing suggested. There were many helper functions of 3-10 lines mapping data and returning back to a single place in a file.
I'm not that smart. Rather, I've done a lot of academic, clever, and sometimes outright dumb things, I haven't forgotten those lessons learned.
On the other hand, it greatly simplifies the process the process of understanding what a function is doing at a high level. It also makes it much easier to navigate the code quickly and determine exactly where changes need to occur.
If you need a whole program laid out consecutive for you to understand it, that's not great.
Sure, _there are_ circumstances. They're just extremely rare and their occurrence is dependent on the team as well. So generally, I'd say function length a pretty good proxy for complexity--as long as there are overrides.
If you try to not always think about contrived cases while writing code, then this rule can sometimes lead to quite long functions simply from executing steps in sequence ("script").
One can split what is really a long function into several small ones for readability, but if one does so I think one should declare that function within the calling function (in languages that support this), and whether one does that or not is really about as insignificant as tab vs spaces or where one likes to have curly braces.
My point is, if a function only has one callsite, and is not likely to get more (incl. test code), then whether to inline that function or write it seperately is a style/formatting issue, not a maintainability issue.
I disagree. I split out anything that I can describe/test separately. If I need to take a list and
1. Filter it based on custom logic 2. Sort it based on custom logic
Then I will generally break the filter and sort out into their own functions even if they're not used anywhere else. Each function can have it's own description and tests.
As you simply describe your own stylistic preference I am not sure if you agree with me about that or not?
If you actually write a seperate test for the function then you have two callsites and of course it must be a seperate function.
1. Static analysis tools should not generate useless noise. I am already well aware that that function I just edited was 1,000 lines long. If I didn't break it down, it's either because I didn't think doing so would be a good idea, or because I don't perceive that to be a problem in the first place, or because I simply don't care. In any case, there's a 100% chance that I'm going to ignore any advice from the static analysis tool on this one. That means that such a rule is, at best, useless.
2. I'm generally not sure if my feelings about long functions reflect reality. I recently saw a talk about evidence-based software practice where the speaker claimed that nobody's been able to demonstrate a link between average function length and code quality, and not for lack of trying.
I ran the stats retrospectively on a codebase I once worked on to derive sensible values for max class length and max function length. To be honest, they came out at some fairly sensible values and I would be happy putting a blocker in place to say "if you are doing anything more than this, you shouldn't be". I think it was something like 200 lines for a method and 3 or 4 thousand lines for a class.
If you can do it at the outset what you'll find is that the vast majority of the time the rule is never triggered. In the majority of cases that it is triggered, it's trivially easy to break the method up a small amount to be under the limit. It's very, very rare to feel like you're having to break it up unnaturally and it feels very forced, though it happens on the extremely odd occasion. I'm OK with that trade-off. The benefit is you never wind up with the 5k line long methodzilla that someone thinks is a good idea and everyone laments forever.
Every code base I've seen has one godzilla method. I'd prefer we avoid godzilla methods at least. You don't need an 18k line class. Nor a 5k line method.
The point is the static analysis tool fails the build, so it becomes physically impossible for the godzilla method to exist. I trust the machine to produce a predictable result. People are a huge source of entropy and having a little bit of automation to help contain some of that entropy I think is a very wise use of technology in CONJUCTION with promoting a culture of caring about quality and diligent review.
It's hard to imagine how anyone ever lets godzilla methods and god objects and all other kinds of code smells become a thing... and yet... they're a real phenomenon that seem to make their way into large code bases at some point or other.
I've worked with people who write messy, tangled code in giant methods, and my experience has been that forcing them to write smaller functions is futile. Satisfying the static analyzer is just too easy: Move all your variables into mutable fields that all the methods fiddle with, break the contents of the method out into smaller methods with cryptic and misleading names, create a master method that calls them all in a row. Voila, with barely any effort you've satisfied the static analyzer and made the code even less maintainable.
Human factors always mess up software... _sigh_
I agree with other comments that Spaghetti Code is not the right term for this. But, the original point is valid.
The other part of the point is also true. To hide implementation inside functions makes sense if the function is part of the real world domain. To do a cut/paste to extract lines of codes of a big function to make it look nicer may increased complexity as now jumping around the code is needed to read it.
A few hints about the language for the non-functional crowd, then my point
1) Functions can have multiple definitions and are picked top to bottom. The one that runs is the first one matching the arguments in the call.
2) Variables are Capitalized (and can't be assigned again, big surprise.)
3) Atoms (:symbols in other languages) are lower_cased.
4) The arrow (->), the comma (,) and the dot (.) do (more or less) what Java/C do with their { ; } characters (sorry, I have to keep it short)
You can start from the bottom and go through the get_token, get_line, parse_ComponentTypeLists etc, find where they are called and see how they neatly (IMHO) replace the switch statements in other languages.
Is it easier to understand? Hard to tell: if the problem is difficult the code won't be easy. Is it less tangled? I think it is. My point: some languages are messiers than others.
[1] https://github.com/erlang/otp/blob/master/lib/asn1/src/asn1c...
Terminology aside, he does make a good point: if the domain can't truly be broken down into clean parts, encapsulation can do more harm than good as you have to weave all these contingencies through every layer.
>The point I am trying to make in this article is that sometimes it's the problem domain itself that's complex and convoluted. In short, it's a spaghetti domain.
I definitely have encountered spaghetti domains and have bashed my head against them, trying to find nice, clean and elegant abstractions that describe the problem succinctly and clearly.
Sometimes there's just more exceptions than rules, and spaghetti code is not a bad outcome. As with all guidelines, do not follow blindly where a less "conventional" solution makes sense?
It's called spaghetti because it's hard to see what information flows where, not just because some information flows are long (at least, my mental image for spaghetti-code has always been cooked spaghetti).
goto might not literally be used any more, but I see "break" used a lot, and for long functions that have variables defined outside of the scope of the "break," it behaves similarly.
More importantly, author seems to just give up around
> After few such iterations the code that was originally nice and manageable becomes a mess of short functions plagued by lots of behavior-modifying arguments, the clean boundaries between components and layers are slowly dissolved and nobody is sure anymore what effect any particular change may have.
Yes, that is how spaghetti code happen.
But, if all those conditions were embedded in a single 1500 lines long function, the situation would be no better. You would have interconnected monstrous system of conditions.
Just knowing that the line exists makes a big difference!
I worked at a place where we had to interact with financial exchanges, which nominally talk standard protocols but actually all have their own quirks in their implementations of the protocols, which we had to account for on our own end. They ended up going with the copy/paste approach, and, while I was initially horrified, I came to realize that they're right.
The admonishment against copy/paste is meant to deal with the risk of accidentally forgetting to update one of the copies when there's a new change that needs to be pushed across all of them. It turns out that, for the problem domain we were working in, which sounds quite similar to the one TFI describes, that never happens. Even when the standards body releases a new version of the protocol, each vendor upgrades on their own schedule, and the slowest-moving vendor might stick on the old version for 5 or even 10 years, and when they finally do get around to it, their implementation of that feature will have its own quirks, so just setting a flag saying, "Vendor X is now on version Y" wasn't going to work, anyway. So then you (hypothetically) create a set of flags saying, "Vendor X is using version Y, but with these behaviors that differ from everyone else's version of Y", add a bunch of new if-statements, and the actual logic that might execute at run-time gets even more obfuscated behind a tangle of counterfactual configurations that won't happen in real life but still exist in the code and are supported in the application.
So, to sum up: nominally the same protocol, but everyone does it differently, and upgrades on their own schedule. It really is a horrendous tangle. Dealing with that mess by mirroring it in your code with an equally tangled pile of flags and conditionals is also going to be a mess, no matter how you choose to organize the code.
Don't do that. Alexander the Great was right: Sometimes the best way to deal with these sorts of things is by cutting them into pieces. Then, when you're dealing with one implementation's quirks, you only have to be thinking about that one set of quirks, and not the combinatorial explosion of possible interactions that you get when you try to mix them all together.
This then inevitably causes the resulting spaghetti code.
https://en.wikipedia.org/wiki/Spaghetti_code#Ravioli_code
(Wanted to link C2 wiki but it's a goner.)
You cannot fix essential complexity with any kind of design. If there are options in design, your code will reflect them at some layer. You might want to suggest, say, Component or entity-based or other design, or mixins, or other object level patches for conditionals.
Spaghetti if otherwise well structured is quite understandable. The more advanced versions? Not nearly. Spaghetti is explicit at least.
That’s certainly true; my favorite method at the moment for this sort of thing is to make it data-driven: have a structure that contains all of the parameters that change based on vendor (including enums/flags for optional behaviors) and a single process implementation that refers to it when needed.
A joke I sometimes make is that spaghetti is best of a plate or in a box from which it cannot escape. Otherwise it would spill.
Ravioli does not spill but the relationships become implicit.
This is an awesome term and I must admit that I've never heard it before.
It does apply to a lot of code that I've seen, and let's face it, written, over the years.
Spaghetti code is almost always just bad design and coding.
What most people call spaghetti code these days is just code. Useful programs are complex and have complex code. Modern software might be difficult to follow, espcially if you're at an amateur or junior level (which is probably 80% of the people out there reading code, cf. Dunning-Kreuger). It's not generally spaghetti.