Jackson structured programming
en.wikipedia.org
en.wikipedia.org
For those who don't know JSP, I’d point to these ideas as worth knowing:
- There’s a class of programming problem that involves traversing context-free structures can be solved very systematically. HTDP addresses this class, but bases code structure only on input structure; JSP synthesized input and output.
- There are some archetypal problems that, however you code, can't be pushed under the rug—most notably structure clashes—and just recognizing them helps.
- Coroutines (or code transformation) let you structure code more cleanly when you need to read or write more than one structure. It’s why real iterators (with yield), which offer a limited form of this, are (in my view) better than Java-style iterators with a next method.
- The idea of viewing a system as a collection of asynchronous processes (Ch. 11 in the JSP book, which later became JSD) with a long-running process for each real-world entity. This was a notable contrast to OOP, and led to a strategy (seeing a resurgence with event storming for DDD) that began with events rather than objects.
[0] https://groups.csail.mit.edu/sdg/pubs/2009/hoare-jsp-3-29-09...
If I remember correctly did the book clearly point out backtracking as a standard method, while mentioning that most languages lacked that, so it had to be implemented manually.
1. Analyse the problem and derive appropriate data representations, write out illustrative examples
2. Write a signature for the function - what does it consume and what does it produce. Spend time to get a concise definition of the computation it performs.
3. Create some illustrative examples of what the function does
4. Outline the function
5. Fill in the gaps to complete the definition of the function
6. Tidy up by converting some of those illustrative examples into tests
There’s a further practice to this of iterative refinement - taking what you learn as you apply this process and use that info to refine what you did.
This is a pretty solid way to go about things and it’s utterly trivial to bring a colleague up to speed with what you’re doing at any step when you need a hand. The biggest gap in practice of this method that i’ve found is a lack of concern for efficiency. It’d be fairly trivial to produce an n^3 solution following these steps when an n solution exists.
I was always swayed by the “make it work, make it right, then make it fast” approach before.
(1) … Actually i just went and opened the book, the actual reference is to Michael Jackson’s method for creating COBOL programs which is the progenitor of JSP.
Structured Programming was thought of as revolutionary. Most folks were either doing COBOL or Assembly, at the time. C was just beginning to feel its oats (It was still thought of as a mostly academic language, but it spawned a few languages that were considered "workhorse" languages, like PL/1).
I did start using Pascal, in the 1980s, because that was Apple's native language. It was a very strange language, coming from Assembly, FORTRAN, BASIC, and PL/1.
I wrote a tool or two, in PL/1.
I think I've come back around to seeing some merit in the "traditional" version. All the state is declared before the main loop. You could stop the program after any iteration of the single loop, save the explicitly declared state, and restart where you left off. The logic inside the loop can be applied to any record and state.
In the double-loop version, your place in the control flow is part of the state. If you stopped, you wouldn't know where to start up again. There's also a state variable (firstLineOfGroup) declared inside the loop.
I know why I prefer the "traditional" version: it can easily be refactored to be functional, and making the state explicit means it doesn't have to be refactored if I want to store the state externally or switch from a batch architecture to a streaming architecture. The JSP version is inherently imperative and needs to be refactored before it can be used in a different architecture. Funny how ideas like that alter your taste.
"So how many of you are familiar with the work of Michael Jackson" Most of the crowd raises their hand to which he replies "Really? Quite a few! Well as you know, Michael Jackson developed Jackson Structural Programming".
He was seemingly unaware of the popstar of the same name.
Maybe I'm only just discovering for myself the "power of typed languages" (feel free to rib on duck typed devs here) but it feels like a huge progress in terms of productivity, bug count and readability. I love modern python.
I was once sent on a work course to learn how to write COBOL using Jackson Structured Programming. I _loathed_ it. I remember thinking that
a) it tried reduce the programmers role to little more than that of an automoton b) it was completely at odds with my views (30 years ago and still today) that good software development is a blend of the technical and creative/artistic.
Happily very little of my working life involved actually using JSP (well Java Server Pages excluded).
At one time Dijkstra had to actually argue for structured programming. No-one would take the counter position today.
The equivalent debate nowadays is between languages that separate statements and expressions and expression-oriented[0] languages.
[0]: https://en.wikipedia.org/wiki/Expression-oriented_programmin...
I mean, how many libraries with weird APIs were created just because you can't write the following in most Algol-like languages:
return if(sth) {
foo = SomeProcessing();
Transform(foo);
} else if(sthelse) {
Transform(someDefault);
} else {
someErrorDefault;
};
In many languages, people are resorting to "immediately invoked function expressions" to simulate this pattern (whether for returning or assigning). And those that can't, well, here goes another pointless little function to encapsulate it[0].--
[0] - Which transitions into what is a real debate - "lots of small functions" vs. "fewer but larger functions". It's one of those holy wars that can't die, because both sides have good arguments. But it's a false choice - a limitation of the tools we're using to write programs. 'emilprogviz has a nice summary of that last point here: https://emilprogviz.com/ep05/ep05-transcript.html
Linking to transcript, in the spirit of "text with screenshots is almost always better than a video" - (nice job providing it Emil!) - but the video itself is good too, as are others on that site.
To offer one argument: the problem with expression oriented languages is their generality, in that you can write "weird" expressions, e.g.:
let x =
[giant block of code with multiple nested lets, ifs, etc.]
in
f(x)
I've definitely done this a number of times. Languages that separate statements and expressions force you to break things down further and prevent the code from going too far to the right.Also, for low level languages, the mapping between code and assembly is clearer in languages that separate statements and expressions (I need a more succinct term).
EDIT: Also, by preventing statements from being used as expressions, you encourage breaking up long and complex statements into smaller functions.
Oh, I see what you mean. I discovered Lisp after working in C++ and Java, so I avoided this pitfall in my code, but if I had a dollar for every instance of:
(let ((some-variable (progn
(do stuff)
(do other stuff)
(let ((some-helper-var ...))
(some more code with some-helper-var)
(more-of-the-same (progn
...)))
;; 50-100 lines later, just return one of the values
some-variable)
that I saw in Lisp codebases, particularly in Emacs, well... I could feed my family for a month or two from that money.There's plenty of abuse potential here (and for the love of god, if your language has 'let', it probably also has lambdas or 'flet', use that to create local functions...). But mitigating this, arguably, is a problem for style guides - the overall feature of "everything is an expression" is powerful. It has nice simplicity to it, and reduces boilerplate :).
> Also, for low level languages, the mapping between code and assembly is clearer in languages that separate statements and expressions
I'm guessing this is where the separation originally came from. Assembly is essentially statement-only. But at this point, I think all programming languages in use crossed the threshold where we're actually programming to an "abstract machine", and being expression-oriented seems to confer greater expressiveness.
> (I need a more succinct term).
"Separatist"? :).
> EDIT: Also, by preventing statements from being used as expressions, you encourage breaking up long and complex statements into smaller functions.
Ah yes. The real debate. I edited my comment to mention it before I saw your edit :).
The simplifying potential of expression-oriented languages is huge. Alan Perlis said[0]: "symmetry is a complexity-reducing concept; seek it everywhere", and this is a great example of that.
For example, languages like C and Ada have both if statements and if expressions, this is a duplication that can be eliminated by making them expression-oriented so you only have an if expression like in Haskell or ML.
But, interestingly, there is one historical case of a language going from expression-oriented to statement/expression separation: ALGOL-W[1] was expression oriented, Pascal[2], its successor, separates statements from expressions. Wirth designed both.
I don't know what the motivation was, but I suspect it's because Pascal was designed to be an educational language, and Wirth must have thought that separating expressions and statements made didactic sense when teaching programming as a recipe or list of things to do, as opposed to the more mathematized formulation of expression-oriented languages (of having an evaluation function from expressions to values).
The successors of Pascal (Ada, Modula and its sequels) retain the statement/expression separation.
[0]: http://pu.inf.uni-tuebingen.de/users/klaeren/epigrams.html
[1]: https://en.wikipedia.org/wiki/ALGOL_W
[2]: https://en.wikipedia.org/wiki/Pascal_(programming_language)
And the block expressions in Algol-68 were certainly used and abused. Basically, the type and value of a block (BEGIN ... END or ( ... )) are those of its last expression, so you can put them anywhere. For example, Algol-68 has the looping construct WHILE <condition> DO <body> OD, but not C's `do <body> while (<condition>)`. So what do you do if you want to have the test at the end of the loop body? Simply
WHILE (
<body>;
<condition>
)
DO SKIP OD
I guess you'd get used to those idioms eventually, or some coding conventions would have arisen if the language had been successful. But I can understand language creators looking at that and seeing how getting rid of it simplifies not only their compilers but also the programs written in their languages. test := True;
WHILE test
DO
<body>;
test = <condition>
ODThis is one of those anti-patterns I sometimes find myself falling into (I don't code Lisp or its descendants often). It's kind of discouraging, to be honest, because it feels like there ought to be some more efficient way to do this, and my inner critic comes along and complains that I'd be better of writing Python or C than learning to do it the Right Way in Lisp.
Is there some kind of idiomatic way to avoid this and write cleaner code? Is the solution to simply extract the progn into another function? Does that violate some rules of function encapsulation in Lisp?
The problem is that the toplevel of a module is full of functions at various levels of granularity.
It might be nice if programming languages had a concept of "code sections" (you could implement this with literate programming), where modules are organized hierarchically into sections, and declarations can be public or private within a section. So you might have:
section foo
// Accessible from outside this section
public important_function()
// Only accessible from inside this section
private utility_function_1()
private utility_function_2()
end sectionExactly--this is what I was hinting at with "rules of function encapsulation." I took a course on LISP in college (apropos of nothing, it featured a lab called, "Isn't this just a one credit course?") in which it seemed like the LISP Way was to use helper functions. It's always felt like something was missing in my ability to "translate" between paradigms because this rubs me the wrong way (although as I went through some examples I realized I do this with some regularity in other languages--but it doesn't "feel" as bad).
In retrospect probably part of the problem was that we ALGOL-adjacent undergrads didn't have the scaffolding to understand more generic approaches like fold and cousins. It was a different way of thinking.
I like the idea of literate programming as a potential solution, especially as it can be used to guide newbies through the code and identify areas where idioms are much different between one's existing paradigms and that of the codebase.
The arguments against it are always that it requires having keywords that behave differently when being expressions or statements.
Often the alternatives and compromises proposed always have the same issues, such as Javascript repurposing the do keyword for enclosing expression ifs.
For instance with Imperial vs Metric System, Metric optimizes for multiplication, while Imperial optimizes for division. You can never unify these under the current system. But you can, if you change our base to base 12. Then they suddenly merge.
With CLI vs GUI, we've realized that we needed a mixture. We need a GUI that runs through a CLI. And now we have that, it's called a website. I think tabs vs spaces was solved similarly, with tabs as spaces that editors config can treat n-straight-spaces as tabs.
I'm firmly on the expression only and lots of functions sides, but that transcript is very interesting. You can already do local functions in C#, it seems to be implying that we should strive for that to be our mixture.
I strongly agree. Also, that's poetically put. I'm saving this in my quotes file.
I like to think that there is no OOP vs FP, there is no strong typing vs no typing, there should be just adding to a shelf with tools that each has application in some situations but also constraints on when it can be used effectively. And your job as developer is to understand the bounds on application and effectiveness.
The road to mastery of development then should be by learning and understand those various tools (in the broadest possible sense) rather than by forming strong opinions and shunning the other side of the debate. People who cut off themselves from OOP will never learn its benefits just as people who cut themselves from FP.
I think you're overstating the case with
return if(sth) {
foo = SomeProcessing();
Transform(foo);
} else if(sthelse) {
Transform(someDefault);
} else {
someErrorDefault;
};
You can transform this to the following in even the most limited ALGOL-like languages, which is less clear but not nearly as heinous as the alternatives you mention: var rv: SomeType; -- return value
rv = someErrorDefault;
if(sth) {
foo = SomeProcessing();
rv = Transform(foo);
} else if(sthelse) {
rv = Transform(someDefault);
}
return rv;
The more popular ALGOL-derived programming languages like C, Java, and post-walrus Python have conditional expressions and assignments inside expressions, which means that in this case you don't need to resort to declaring a variable. In Golang and Pascal you can just assign to the named return value rather than declaring it as a normal variable and then explicitly returning it.The argument against separating expressions from statements is also that it adds redundancy to your programs, as in the above example.
I wrote an essay about this tradeoff in general, not limited to expressions vs. statements, at http://www.paulgraham.com/redund.html.
Imagine if, nowadays, a grad student took a node in a scientific cluster to run Dreamweaver in a Windows VM instead of writing HTML by hand. (Sorry, it's the closest analogy I could come up with)
He was probably frustrated that people couldn't see the machine code fully formed in their mind's eye :)
I don't see expression-oriented as being similar to that at all. Expression-oriented may provide a boost to programmer efficiency, but it's not nearly on the order of structured programming vs. unstructured. It doesn't require changing peoples' mindset to realizing that programmer time is more valuable than machine time; they already know that. And, the parallel problem would be clarity of data flow, and I'm not sure that expression-oriented is a huge win in that area.
In my opinion it is, because you can compose data transformations from smaller parts. Each part can do only one thing.
For example in Clojure, you could do.
(def transform
(comp
(do-thing-a)
(do-thing-b)
(do-thing-c))
(map transform my-collection)
Or you could use a threading macro, with thread first or thread last semantics, which can also help you build clear pipelines. (-> basket-of-apples
(select-ripe)
(clean)
(cut-to-pieces-of-size 4)
(pack))
Not to mention transducers, which can further improve and clarify data flows.So I learned the usual lesson many do when you go from Educational ideals into Working reality. Real World business don't run on cutting edge changing soon as that changes, legacy/stability and historical aspects do play out. So whilst you may know a better way, there are many factors that make that impracticable. Sure if rewriting your code-base and hardware to the cutting edge of the time was viable then people would do it, but testing and verification of code - when done properly takes longer than the time to build the latest cutting edge system. Let alone the whole cost factor.
I imagine many have comparable stories upon their move from the realms of Education into Work. Do share as be nice to read what brain-walls today's first workers encounter.
JSP's claim to greater clarity is based on the correspondence between the code blocks and the structure being processed: the body of the outer loop processes a group (of gap and word); the body of the inner loops process a gap and a word respectively. The traditional style, in contrast, is like a state machine, and can't be viewed structurally---you need to think about what's happening at each point. This makes it harder to modify the code (eg, by adding something that happens once per word, which is now in two places) and often leads to bugs.
This same structural clarity is why, in my view, code that uses list functionals like map/reduce is usually more comprehensible than traditional imperative code.
// JSP version of a split function
split_jsp = function (s) {
words = [];
i = 0;
while (i < s.length) {
while (is_white (s[i]))
i++;
word = "";
while (is_alpha (s[i]))
word += s[i++];
if (word.length > 0) words.push(word);
}
return words;
}
// "traditional" version of a split function
split_traditional = function (s) {
words = [];
word = "";
for (i = 0; i < s.length; i++) {
ch = s[i];
if (is_white (ch)) {
if (word != "") words.push (word)
word = "";
}
else
word += ch;
}
if (word != "")
words.push (word);
return words;
}
is_white = ch => (ch == ' ' || ch == '\t' || ch == '\n')
is_alpha = ch => ch != undefined && (/[a-zA-Z]/).test(ch)I worked at Eastern Electricity Board in the early 80's - No JSP there, but did work upon a new project that was object focused - using COBOL - no data dictionary overkill project, but saw what would normally be one program broken down into a main body that would link in lots of functions - which in themselves would be small contained COBOL programs. So one data input screen would see a program for each field type, instead of one large program and the main code would in effect be a skeleton that would link in the screen template and needed programs to handle those input/output data fields. Certainly the way forward in many ways, though COBOL perhaps not the most learning towards that.
I liked JSP, but then that was how I was taught COBOL in education. Never really got to use it in anger due to legacy standards and other factors which alas made sense unless your doing greenfeild at the time.