The Wrong Abstraction (2016)
sandimetz.com
sandimetz.com
IMHO abstraction should not be guided by the desire to remove duplication. Duplication is not even the only (and far from the worst) result of insufficient abstraction.
Insufficient abstraction leads to increased complexity, not just duplication.
Example: just this week I've been working on some code that has to deal with arbitrary ranges of ordered values. Typically when you think of a range, you think of a pair of bounds - the lower and the upper bound. However, the input is allowed to have only half-ranges so that one of the ends might be unbounded. So in the code I inherited there are 3 cases: a range with both lower and upper bounds defined, a range with only a lower bound, and a range with only an upper bound. All code processing those ranges has to deal with that optionality of either end, thus making it way more complex than needed - lot of if ladders or switch statements. And it multiplies very quickly when you deal with more than one range at a time. It is insufficiently abstract, even though it doesn't have any obvious duplication. The proper abstraction would be to transform the half-ranges to full ranges by introducing special open-end items (always smaller or greater than every possible value) which would allow one simple type of range to cover all possible cases.
The post already says: −∞ < x < ∞ for all numerical x. (And the mathematician in me clarifies that that's all real numerical x.)
Or, you can use a range which is a sum of three/four cases, and not worry about any of that.
> If it is over the reals, sure, you can use -∞ and ∞. If it is over the integers, you can use MIN_VALUE and MAX_VALUE
If you already knew the answer, why did you ask?
> What are these special open-ended items? How do you need to extend comparison to account for them?
The whole point of abstraction is to make those decisions once and isolate the complexity in one place instead of having it spread over N places in the code, forcing everybody to solve the same problems again and again.
I would think of the range itself as the abstraction, and then it matters less how it's implemented since any potential problems are local to the implementation and cheap-ish to fix.
Encapsulation of ranges doesn't actually fully solve the problem, but just moves and isolates the complexity to the private implementation of the range concept (likely a class in OOP). E.g. you want to compute if two ranges overlap - you still have to deal with the complexity of 3 cases in each argument, so total 9 cases.
And hiding the bounds is likely going to be a lot more intrusive on the existing caller code (more refactoring).
Or could use both. ;)
The point Sandi Metz is making is that it's often only obvious when you can do this in hindsight.
Lots of things can look abstractions worthy but aren't - including things that are repeated 3 or even 10 times.
Intervals are nice for representing acceptable ranges. Half intervals mean greater/less than. If you stick infinities on the ends, everything likely works. You then expose methods or functions for all your operations. From the outside, you don’t have to care if it’s a half interval or not (unless that is what you’re particularly checking). On the inside you don’t really, either.
If you’re messing with intervals in a business setting, it’s worth considering if you need multi intervals, non continuous regions.
These are all great for handling uncertainty. Like if you add two weights that have +/- values, you can have the sum have those and be correct. The math is all well defined and rather easy. Wikipedia has good pages on it.
You wouldn't even need to create anything new—both math and C already provide this abstraction in the form of −∞/-inf and ∞/inf.
Both insufficient and wrong abstractions are viral. They infect everything they touch, which can snowball into large parts being more complex, harder to understand and debug and often also slower.
The wrong abstraction is wrong, insufficient abstraction is wrong.
Really the only weapons against complexity we have as programmers are decomposition and abstraction. We have to take things apart, like in your example it would be the meaning of each parameter, and then we put them together in such away that the details below our abstraction can mostly be ignored.
I say that all with a caveat: I tend to prefer less, insufficient or no abstraction over the wrong one. The former few options can lead to code that is hard to understand as a whole and can be brittle, but the latter drives you into a corner: The only way out is either trying to patch over it or starting from scratch - choose your poison...
Often enough, the way out is going back. Why are developers collectively so reluctant to go back? (Myself included.)
Those concepts are so alien that I believe if I say those words on a recent conversation, lots of people will pop trying insisting on redefining them into meaningless ones.
In the example you set, it's the right time to apply an abstraction, so it's no longer premature. Perhaps the maxim should be labelled as "premature abstraction", rather than "premature optimisation".
I'm using Tonal which makes it easier, because I can mostly push weirdness into wrappers for individual Tonal calls. It's honestly been a great little challenge because the scope is so small that it doesn't take all that much analysis or thought to see where abstractions break down. Fun little exercise in code design.
I mean, at a high level, theory is the abstraction isn't it?
Music theory in general is a somewhat difficult abstraction due to the multiple ways to interpret different things in different contexts. The same chord progression might be thought of as being in several different keys based on other contextual information for example.
I like this phrasing a lot, thanks for this!
I'm still wondering if there's also potential in avoiding the wrong abstractions in the first place. For that we'd need a "cheap" way to decide whether an abstraction is good/bad/something else.
Is there generally applicable, widely accepted principles or research around this? A quick search only revealed random blog posts; nothing I'd consider widely accepted.
J. Ousterhout is gaining traction, at least in my corner of the industry. https://web.stanford.edu/~ouster/cgi-bin/cs190-winter18/lect...
Unfortunately, I feel the the original problem can be obfuscated by adding ideas to the existing problem who now need to understand your ideas (or mental model) of the problem to understand your code. I need to understand how you think to read your code. And if your way of thinking is more advanced than mine or incomplete or not great, then my work is harder.
Mental models are how I understand software, there is what I call a "critical insight" that makes the code obvious and easy to understand. I don't want to be deciphering and spend days investigating code to understand how to change it or build upon it or use it. I want the APIs to reflect their expected usage and behaviours.
My perspective is that computers are adding and arrangement machines - they add and do operation on numbers and move things around to different locations. My mental model of computers is that it comes down to LOGISTICS and arrangement/ordering problems. Unfortunately, APIs and data structures are nested and ordered and obfuscate the underlying movement of things between places or addition to different things.
Everything that obfuscates the rules of the computation means getting the behaviour you want from the code is harder.
I've often thought about "commutative computation" where we specify what we want to be true to the computer and the computer works out how to arrange all its existing computations to satisfy that additional invariant. I often think of software as a series of behaviours rather than functional or imperative.
Think of a materialised view, we have an existing behaviour of the computer and we want to customise the behaviour. You could work out where you need to insert your code snippet into but that's really hard. Or you could add an invariant to the system that the system now satisfies.
What do you think are the steps someone goes through when they obfuscate the problem by adding ideas? Like why do you think they do it?
I try think of the simplest most elegant, beautiful solution to the problem that allows the minimal of code and minimal cleverness and complexity be used to solve the problem with trivial loop, map, hash lookup or traversal or association.
That usually comes with trying to see the problem in a different light, to reframe the problem as a different kind of problem, which can obfuscate the original problem.
Eg. Iterate through an array:
const arr = [1, 2, 3];
for (let i = 0, l = arr.length; i < l; ++i) { console.log(arr[i]) }
Let's model it differently using an iterator: const arr = [1, 2, 3];
const arrIter = arr[Symbol.iterator]();
let i = arrIter.next();
while (!i.done) {
console.log(i.value);
i = arrIter.next();
}
At this level it's still pretty obvious what's going on, but you can still see that there's a level of abstraction between an array access vs calling 'next/value', and that obfuscates what is actually happening at the computation/instruction level.If I extend this another level then I'm going to start modelling problems using an iterable and not an array/index. New requirements come in and we extend to use an async iterable. Everything still works nicely, but in some scenarios where the actual iterable is just an array, now there's a lot of extra overhead to just do an index lookup.
Using the iterator allows the code to be reused in more scenarios, but there's usually a cost to switching the lens of abstraction so that it fits into a problems modeled differently.
The other thing is that unions/intersections are not an abstraction, because they don’t hide any details. The purpose of an abstraction is to separate essential properties of whatever is being modeled (the interface) from current details that may change later, or that client code shouldn’t depend on (the implementation).
Conversely, I am currently working with a frontend code base that is using "classic CSS", and it's striking to me how frustrating it is to have to think up what the "semantics" of this and that particular <div> can be said to be, when there very often aren't any.
1. Organizations must value continuous improvement. We want to avoid two extremal behaviors that sours individuals. First, the lethargic in-bred sterility of: hey, it worked before you got here, and it's fine now. Play-it-off is not wisdom. On the other extreme is frustration gone wrong. Sure, you can see a problem AND be right about it. But whining and constant criticism sours. Everybody's problem is there are 10,000 things that could be worked on, and resources only for 1000. You better make sure you're customer driven so you pick the right 1000.
2. Duplication is better for the medium term ... if you stay with the problem for a while, you are better able to distill the big picture into a more coherent new abstraction. Here you can cite a problem, cite a solution, and stick to your guns. You are better positioned to impact change without being a whiner. Now, problems are working for you, not against.
Fortunately my current project is different, because the team is very small and we have silos of responsibility, so we don't really get in each other's way that much.
It appears that the largest obstacle here is not the lack of ability, but agency.
In my experience, this is the failure point in thinking about incremental development. A portion of code's mere existence does not make it untouchable -- I constantly express this to my team. If we are presented with new information (such as a capability) that would function best with a refactor, so be it.
We have some team members who think the opposite -- that any existing implementation must be preserved (I call this the "incumbency fallacy".) It creates a weird motivation to implement quickly/early, as if getting your pull request in is the most important aspect of your work. And yet, what I find is that early implementations often inform how something should be improved after the fact.
"Re-introduce duplication by inlining the abstracted code back into every caller."
Or maybe if there are, say, 10 places dependent on the shared code, we can make them 5 places dependent on one version and 5 places dependent on another?
Forking and merging is part of business as usual in programming and we should be used to it. We should not be shocked that sometimes you have to fork a function because adding more parameters is not feasible, but nor should we declare sharing is therefore wrong or harmful.
Also, how you design parameters is extremely important. One callback parameter may be worth a hundred "normal" ones.
If midpoints are introduced then comments like yours "but what about..." can always be made until the entire abstraction tower is fully described, and that's not the blog post (or book) the author wanted to write.
Have you ever seen an online argument? If someone is right, and someone is not AS right as they are, they are a "left shill" and vice versa. If you promote solution A, and someone promotes solution B, then they're "wrong". Not establishing points, just "wrong".
So I think establishing two directions is best accomplished not by marking up two extreme points and leaving the rest to the imagination as our imagination is apparently quite poor.
It's more correct to describe the next step in a direction, and let us take things step by step and know that nuance is inherent to our success, not optional.
I’ve probably paraphrased this line to every junior engineer I’ve mentored. It’s such a succinct and pithy insight.
- know your customer, and their use cases
- continuous improvement
- specialists have got to know the big picture and engage in it
- good cross functional coordination
- CS fundamentals: DBs, algorithms, UNIX, functional programming, C/C++, CDCI etc.
That never ages away.
However, stuff like we use and have always used Kafka (read the code!) for messaging, so we're not doing kernel-by-pass to move data now is small 'c' culture.
Small 'c' culture is the kind of stuff that, if you abrogate it, a small army of people will come out of the woodwork and brow beat you for it. Brow beating to keep you inline is not engineering. It's nagging.
Tradition, when it's small 'c', is stifling. Don't fall for it.
https://en.m.wikipedia.org/wiki/Rule_of_three_(computer_prog...
To push back a bit on naive misapplication of DRY I've been saying we should call collapsing things that are just coincidentally similar (and likely to change independently) "Huffman coding".
Combine, reuse, rewrite as needed.
Eliminate control flow.
- Don't fall in the trap of early abstraction...
- and abstracting one single use case is very hard, if not impossible...
- but you are writing your app in Go and interfaces are the only way to properly test your stuff...
- The end.
Adding tests using interfaces is solely adding more code (and the method signatures are simply duplicated) so it's the exact opposite of OP's problem.
Design pattern is real abstraction, because it's about thinking and designing. Abstraction is not related with specific implementation.
So, duplication is fine until you figure out the real Design Pattern to be used.
So i'd collapse this whole article down to 1 bit - "software development is hard"
Isn't that a pretty useless sentence? Of course duplication is cheaper because aren't the higher costs one of the reasons why it's a wrong abstraction?
Reminds of this sketch from Fry and Laurie
>Hugh: Yes but too much is bad for you.
> Stephen: Well of course too much is bad for you, that's what "too much" means you blithering twat. If you had too much water it would be bad for you, wouldn't it? "Too much" precisely means that quantity which is excessive, that's what it means. Could you ever say "too much water is good for you"? I mean if it's too much it's too much. Too much of anything is too much. Obviously. Jesus.
That's like saying don't do the wrong thing.