Waterloo Style
theprogrammersparadox.blogspot.com
theprogrammersparadox.blogspot.com
A few relevant quotes:
"Bad programmers worry about the code. Good programmers worry about data structures and their relationships." -- Linus Torvalds
"Show me your flowchart and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won't usually need your flowcharts; they’ll be obvious." -- Fred Brooks, The Mythical Man Month (1975)
It's no coincidence that ADTs and data-driven design are commonplace today -- I would even argue they are so commonplace that most programmers are not even aware they are stylistic choices.
Many, many folks get caught up chasing complexity for all the wrong reasons. Define your data, stake your boundaries and just write the damn code.
It's okay that someone blogs about an old idea and puts a new name on it. It's still a good thing even if it's obvious to some. The author came to their own understanding of this concept in a particular way and decided to share that way with other people. You don't have to knee-jerk critique them for it.
However, that only further proves your point. Those tables were a mess, and the underlying logic was even worse.
Me neither.
I lot of these ideas seem to come from SICP which originated at MIT, from earlier research at the AI lab and at Xerox PARC.
I worked with a bunch of smarter-than-me UW grads after graduating.
My “how to write large systems” takeaway from that early point in my career was to focus on the interfaces between various parts. What I’d never thought about until now is that is a very data centric viewpoint.
- What system has what data?
- In what shape?
- What shape does the next system need its data in?
- Are the interfaces between these orthogonal? Shallow? Easy to grok? Tight (as opposed to leaky)?
Great article.
>Shallow?
too many deep inheritance typescript libraries being built today I find!
One thing that amazed me while working through HTDP was that almost all my bugs came from not understanding the data types correctly, or messing something up with the process of manual type-checking. Once I understood the data structures I was trying to pass around and compute with, the bugs almost always melted away.
Now I program basically everything in a language with type checking (mostly in TypeScript) with thinking about data types and type definitions as the foundation. I'm amazed to see how 95% of the pain, complexity, and bugs has just melted away.
[1]: https://htdp.org
(Also worked at a waterloo unicorn. Codebase was the cleanest i’ve worked in, though perhaps excessively so, at the cost of product velocity)
Still, having quite a bit of experience under my belt, I don’t think it’s actually going to be uniquely helpful at writing code. As I’ve gained more mastery, I’ve started to think more abstractly about the system as a whole. Sure, data flow is one aspect and you should consider. The mechanical aspects of code are necessary to consider too (eg what makes code maintainable and robust against mistakes). But how all the different pieces cooperate to create a complex system that’s solving your problem, that’s the way to get some real insights. Thinking about the system, you can start to think about how to change requirements rather than just trying to solve within some external constraints. Being able to move seamlessly up and down the abstraction stack is hugely important.
I agree 100% that focusing on the code is completely misguided. But so is focusing on data. Data by itself is useless. It’s what you can do with the data and how you can use it. Just shuttling bits around is by itself pointless unless all you’re building is basic data viz. And ultimately, this by itself is only one approach. For example, AI systems depend on data cleaning today. That’s not at all about how you shuttle data around not will that perspective help you. ML systems depend on more scientific rigor approaches. A data perspective might help you optimize the performance of those systems, but that’s a smaller aspect of what AI systems are trying to solve (not unimportant, but smaller than the entire thing itself). Smaller perspectives aren’t bad but they limit the space you can play in (which may be the goal some times but keep that in mind).
All that being said, a system’s level perspective is also limiting. You’re some times not going to have the domain expertise to actually solve some problems by yourself. You want to take on lots of different perspectives and have a good sense of which situation falls for a given perspective more. And some times, you may not have the ability to take a certain perspective. That’s where colleagues can help to complement your weaknesses.
Now that is Waterloo Style.
I used it professionally for years in the 1980s. It wasn't just a teaching language.
It turns out doing dataflow analysis is just pretty much scorned by the programming community so it was moot.
People just want to start coding and get that immediate dopamine hit of positive feedback. The answer to which formalism was better was "Agile" where you don't need to plan or even understand data flow (because it will emerge spontaneously); just write code between Ritalin hits.
One result of this is that there are no good (free) tools out there to support either dataflow analysis formalism. I get blank looks for coworkers when I ask to see the dataflow analysis for their systems.
Of course, I picked up the dataflow analysis thing working at a Waterloo startup as my first post-graduation job. The university I went to focused on data structures and algorithms (the how of software) rather than dataflow (the what of software). My first job taught me that data structures and algorithms are necessary but not sufficient.
I like the idea of having models and enforcing themselves. For example, testing that three different API endpoints of a service match each other's idea of their objects. This is a sanity check when we verify the frontend state.
If we could separate scraping from modeling constraints, we could potentially collect data separately from the verification step. Then we aren't left waiting for UI DOM stuff when we verify the model. The latter can happen separately, and extremely quickly.
See also
Data-Oriented Programming. Sharvit, Yehonathan. Manning 2022.
Principles of Program Design. M.A. Jackson. 1975.
SAM Pattern. Jean-Jacques Dubray. https://sam.js.org/
He told me, paraphrased since this was 20 years ago, "Whenever I'm designing a new feature, I always look at what needs to come out at the end, then I can figure out what needs to go in at the start and how it has to flow through."
This still sticks with me and serves me well.
I will say though, that I've also noticed the contrast before with MIT grads, who tend to have a very strong LISP bent to their styles. It's true that each school has their own unique flavour, and much like accents it may just be that you don't notice your own.
Isn't this very data-centric in nature?
0. There is no magic anywhere. Anywhere down a stack, in a system or in code, they are all just bits of code. The behavior lives somewhere. (Genchi genbutsu.)
1. Get the point A. to the point B is dataflow analysis. One can even deconstruct the RTL micro- and macrocode designs of a CPU or GPU this way. Input, process, output, and feedback encapsulate represented behavior be it a shell pipe, streaming IO class, Kafka, firewall, audio effect generator, or microcontroller.
Static compilers try to be efficient data flow analysts with as much liveness and constraint information as possible to apply optimization transformations. It's interesting that static optimization passes act are usually implemented as middleware patterns that stack.
I find it tricky keeping 'data-1st' code my frontend (which is eating more & more of backend as the years pass)
My question is how do they fare when they are to use imperative languages later on in the CS program where they have to use messy for loops and mutation and memory allocation? Is it better because they did CS135 first or hard?
To be frank, I don't think imperative language use is going away anytime soon. So, they need to learn the best use of both the worlds, hence, asking.
Unis have always been a bit out of the way, they are not 'sponsored by banks' they were sponsored by Churches and congregations, then the elite.
Waterloo was Waterloo College, a Lutheran Seminary, and grew out from there.
It was successful probably because it was very much focused on Tech, unlike most other schools, and didn't appeal to multi generational families, but 'anyone'. The local mennonites are also extremely good students, you don't hear about them, but they get good grades.
It's a great tech school, but one of the ugliesst, most sparse and uninspiring campuses imaginable. If we think of traditional Uni like an 'Ivy Campus' or 'Oxford' aesthetically - UW is like one of those 1960's, concrete block kind of Soviet Utilitarian places. I mean it could be worse.
The way I came to appreciate the aesthetic of UW was to realize it is a school of the 20th and 21st centuries, whereas the Ivies are schools of the 18th and 19th centuries.
When you think of an abstract university setting, ‘the future’ looks like UW.
During the mid 20th century there was a lot of that going on.
Most 'modern' designs are rubbish, but at least what they make today is much more interesting, for the most part.
In the 1970s, students at Cal State LA called it Cal Concrete.
At this point it’s a reinforcing flywheel just like it is with MIT, Stanford, and Berkeley. I think they went with a different route though. They give minimal scholarships and afaik they don’t go out of their way to recruit wunderkids.
I suspect there isn’t any single answer / magic secret. They just built a good culture centered around teaching kids STEM effectively, kids and parents recognized it quickly enough which created natural competition to get in until it became a flywheel effect.
Originally this was only for professors (and I'm not so sure there weren't conflicts of interest). My first employer fought the university for the right to his Masters' thesis, and won, establishing the precedent for grad students (and hence my job).
From what alumni told me, undergrads are incentivized to apply everywhere for internships as part of their courses and especially during the off-cycles (winter) for internships when they are effectively the only ones looking. Coop also means someone who can't convince an employer to pay for them won't graduate, so there's a nice selection bias.
The large firm involved in the origin of UW was Electrohome, one of the top consumer electronics manufacturers of its time.
You can have correctly modeled your data, but if you then use bubble sort, your program will suck anyway!
But it is true that while most programming can be seen as trasnforming one data to another, it is the algorithms that do that transformation, such as when transforming a list to a sorted list.
But an algorithm can be packaged into a function so that you only need to know what the function does, not how it does it. Data can not similarly be "encapsulated". Except you could say that getter-methods are a way of encapsulating data.
Gets 90% of the way there and is loosely coupled, easy to test, and scale horizontally.
Modeling an ADT sure, but an air traffic control system by following the data flows? Not likely.