The Pros and Cons of DRY Code
qvault.io
qvault.io
> Every piece of knowledge must have a single, unambiguous, authoritative representation within a system.
It says "knowledge" not "code". This is an important distinction. Removing duplicate code might not be removing duplicate knowledge and having duplicate/similar code doesn't necessarily mean you're duplicating knowledge.
A toy example: if you happen to have 2 products that are the same price you still wouldn't want to combine them into one constant value.
Some code might be similar by happenstance and combining that will cause problems when that happenstance stops being the case.
The exception traces go on for a couple of screens, so I guess the coders feel like their CompSci degrees are justified.
(EDIT: this is mentioned in the article as "localization complexity")
There is a special subgenre of functional programming, where the style is to explicitly not have names
https://en.wikipedia.org/wiki/Tacit_programming (also known as point-free)
if (weShouldDoThis()) {
return doIt();
}
Structurally, this seems very similar to having twenty methods with one conditional block and a delegation to some other method. But the stack trace is half as high, and the testing surface area is greatly improved, because the code in the 'if' clause is pure, or close to it, while the 'doIt()' method might alter shared state.There's a big difference from a robustness standpoint between having half of your functions pure, and half of each function pure.
I'm working in audio DSP as a rule, so I already have to maintain state of 'what the audio is doing related to this set of sample data', and often some things about how controls must move and feed information into the program to relate to a desired change in the audio which itself can be somewhat indirect. Maybe I am just stupid, but when I have to track DRY code I just blow a gasket immediately and can't parse it at all.
I've wondered whether some of that stuff is about taking trivial problems and making them hard to follow for the uninitiated, to protect employment. WHY not repeat yourself, when the problem is repeated tasks? Or more relevantly, why not write a complicated and long sentence, rather than mangle the sentence into a series of more abbreviated footnotes? DRY reads to me like that chain of footnotes. You lose the plot.
Well, I do :)
It's a specific example of my general rule in designing my projects: things that tend to change together should be closer together.
After pulling out the data, the code becomes simpler, more general and more composable. The data structure then often reveals repetition that is not redundancy: If you look at a configuration file, a SQL table or a REST response that has repetition, it becomes clear that it represents different facts (to use your term) that happen to have the same values (Alice and Bob have the same age). This form or repetition is fine, correct and often interesting!
So by the practice of pulling out data from code we get DRY (by the definition of the authors) code.
However there is also also another form of repetition that is a bit more subtle: boilerplate. There, the representation of an algorithm (as code) is drowned in noise that does not convey the intent and the flow of what is happening.
To increase the signal to noise ratio in code that is already(!) DRY but clouded by boilerplate, one needs a way to reduce syntax itself, which can be achieved by meta-programming and DSLs.
Often they've gone into removing idiomatic code as if idiomatic code isn't a presumed piece of knowledge the entire team shares, and seeing that code indicates "everything is normal" while seeing a function indicates "something odd may be happening, spend cycles looking at me".
The other definition is requirements, rather than knowledge or duplication. If two pieces of code are driven by separate requirements, eg, the price of two items is typically based on very different business situations, unless the two items are flavors of the same beverage perhaps, then the fact they share a price is an accident of the current state of the business, not dictated by a common requirement that they both be $1.50.
If both have to calculate sales tax because of the same statute, then don't implement the calculation twice. If a new statute changes that situation, for instance due to a sugar tax, split the code when the new requirement comes in.
DRY stands for “don’t repeat yourself”. that’s its de facto definition, unfortunately... whether it’s useful or not.
Just like REST and OOP and a host of other things that “everyone gets wrong”, you have to consider that if everyone gets it wrong, maybe the messaging isn’t very good in the first place.
In my experience, it just means that an army of ignorant middlemen grabbed hold of a concept and ran too far with it.
OOP / UML / XML / etc. etc. They all actually solve problems very well. But the people who taught those concepts were mostly selling snake-oil under the guise of XML or whatever.
If any particular paradigm becomes useful: the parasite class of marketers / con-men come out and sell fake versions of the strategy. Eventually, the con-men completely corrupt the concept and we just have to move on to the new concepts that haven't been corrupted yet.
[0] https://fsharpforfunandprofit.com/posts/designing-with-types...
A more common example, take an online clothes store. They only sell shirts and trousers (pants - :]). It's a really simple app. Obviously, I'm skimming a great deal here, but each product category is described by the following properties:
- Shirts -- Price -- Size -- Colour
- Trousers (Pants - :]) -- Price -- Size -- Colour -- Length
What the DRY principle generally addresses, is this basic idea. My instinct would be to create an abstract class called Clothing:
- Clothing -- Price -- Size -- Colour
Trousers (Pants - :]), for me, would definitely be an extension of Clothing. But until my suppliers can furnish me with Dresses, T-shirts, Scarves, etc, I would not be too concerned about even creating a separate Shirt class, but probably would.
What I've since found through experience however, is that not everyone thinks this way. For some people, each product category would have absolutely have to have it's own wholly encapsulated class in this situation. I can see the problem such an approach attempts to head off, but until the problem actually exists, I believe this pattern makes the development process painfully inefficient, and harder to maintain.
It's interesting to me, because while it seems like 'one of those' arguments, it's somewhat more impactful than 'tabs vs spaces'.
In your example, the store adds "Shoes". Then later, it adds "Sale Items". Now you want some shoes, trousers, and shirts to be on sale.
So each item now has two edges, its "primary_category", and "sale_items". A little bit later, the store turns into a Sale Items at Massive Discounts! Store, and the primary ontology becomes price tiers.
etc
- Clothing -- Price -- Size -- Colour -- Discount
Then does your scenario really become as big a problem as you suggest? For one, some ontological relationships needn't be described in the same layer, or can simply be implied.
A Red Fox, a Red Ant and a Red Snapper are definitely related in some way, but I wouldn't use the same class to program a representation of all three.
A Red Shirt, and a Red Pair of Trousers/Pants though, probably share enough for a rough outline for online shopping app though.
This way I would never need to make a class "shirt" or "pant" or whatever. I would simply add categories for them. Way simpler than encoding this stuff in some class hierarchy. There is that word again, hierarchy. The problem with those is always, that sometimes you will have an item, which fits into 2 sibling categoies or into 2 parts of the tree, which are not on the same path. With classes you will need multi inheritance or you will need to come up with a new parent of both of the classes.
The problem with this approach (and I've done it) is that product-attributes is now this black hole that's difficult to query, difficult to modify, a huge potential for errors that could even be caused by a typo. It's also incredibly difficult to version.
I regret going with this approach even though it results in far fewer tables and classes. It's short term gain for long term pain. Everything is trade off -- it might well be more than worth it for something not expected to live long or have a lot of changes.
Although I agree with your point that not everything needs to be hierarchy. You can be explicit with your data without having everything inherit from everything else.
I guess the approach stands or falls with developer discipline. If you one is not disciplined enough about the conventions and carefully choosing them, one could end up with a mess of a mass of unknown keys in products, which no one knows why or how they got there.
Think tags on items represented on a tree, where clicking the tag reorients the category around what you clicked on.
I could make the tables flexible (add columns to fit all data models). I could make the queries dynamic (make the table part flexible). But I'm duplicating code instead.
As the proverb goes, a little duplication is better than the wrong abstraction. There's bits of code in here where the author tried to be smart, but queries like "select id from $table where x = $y" really throw my editor (and brain) off.
Yep. Though the same pathologies exist even with this framing. By that I mean it's easy to misdiagnose what the "knowledge" actually is. Even totally identical types and behaviors can be distinct in the knowledge sense if they exist in different contexts.
For example: How to render some part of a website for example or a calculation, which uses the same formula as is used in another part of the code to calculate a number.
So with that interpretation of knowledge, you would have to avoid duplicating code. And it makes sense, because, if one day you need to change the way the computer does something, you want to avoid changing it in multiple places, because that makes it easy to forget places. This way avoiding duplicated code increases maintainability.
Yes, he found an abstraction that turned out to be not useful. But that doesn't mean there wasn't an abstraction that could have been useful. Or even that the judgment call that his exact abstraction wasn't useful was correct.
DRY is important. Maybe he didn't find the right abstraction, but for the most part it's possible to refactor repetitious code with readable abstractions.
> but for the most part it's possible to refactor repetitious code with readable abstractions.
The operative phrase being "for the most part".
It wasn't solving any problems.
It's better to focus on solving problems as they rise up, as opposed to doing things to keep in line with some arbitrary standards that may not be relevant.
Key is learn how to discern between trivial issues and real issues that affect the project.
The point of his article was that he'd CREATED possibly significant problems through imposing an abstraction where it didn't fit. I can think of a number of new 'things' for 'resizing' that would be valid new 'things' with intitutively sensible behaviors that would violate any enclosing abstraction he might have produced, especially the one he offered as a 'bad example' where sides are conflated with corners and so on. He's right to have realized he was on the wrong path.
https://angular.io/guide/styleguide#t-dry-try-to-be-dry
> Do be DRY (Don't Repeat Yourself).
> Avoid being so DRY that you sacrifice readability.
In the unit-testing world, they say "DAMP, not DRY". Descriptive and Meaningful Phrases, which happens to be repeating things a lot more than twice!
Some things are just clearer when repeated. Under such conditions, repeat yourself all the time (as long as it is clearer).
The question of DAMP vs DRY is eternal. Repetition is sometimes more clear, but sometimes is overly verbose and unclear. Its hard to know without experience.
What it led to was an inscrutable codebase where unless you had been working with it for years, finding out where a thing was actually done was an exercise not unlike an archaeological dig.
Like most things, it’s a matter of something which is a good idea, but the reasons it is a good idea are lost in its most common application which is at the exteme.
I’m discussing a refactor with a DRY engineer, and they say something like “So what do you suggest? Just replacing that name in - every - single - place - it’s used? That would be error prone and time consuming.” Then I run a find/replace, the unit tests run, the green light flashes, heads explode.
Again: DRY vs DAMP is often a very good thing to be debating, but in my personal experience the DRY advocates have gotten way too dogmatic.
Its super fast & easy to take a business service like CustomerLoginService and copypasta it into a new version called CustomerLoginService2 which addresses 99% of the same stuff but has a small tweaks throughout to help with some special business edge case or migration to a new way of doing things. Adding the new service to DI and being able to access it (in a well-designed application) is typically 1 line of code on both counts.
I don't worry too much about this kind of stuff anymore. If we end up with 5+ copies of this thing and it actually causes problems or otherwise slows us down, then I would probably consider taking the time to actually clean it up using something like parameters, inheritance, or generics to establish desired context of usage and normalize code.
This sort of scheme might look tacky and trigger developers a little bit, but seeing CustomerLoginService2 in a stack trace and being able to instantly zoom to a scoped implementation unique for that context of usage is a pretty big advantage. Contrast with if you were to normalize down to a single CustomerLoginService. Debugging the added conditionals & state will certainly be more cumbersome.
I, in most cases, disagree with this. If this were true, I suspect the best and easiest to understand functions would be thousands of lines long
Often you write functions that only get called from one place, just so you can give a name to a block of code and make it easier to understand.
Just today, for example, I wrote a big case/switch statement to figure out the URL for a user on a system where the URL is different depending on where the user exists in a hierarchy (I didn’t write the server side of this - gotta work with what I have). I extracted this all out into a “getUserUrl()” function. In the caller, it’s much easier to understand what is going on when you see “getUserUrl()” than it would have been if there was the giant switch statement there.
Cognitive burden would generally not be improved by inclining every function I call.
My thought is special case arguments are fine for internal calls but are lousy for external ones.
If you have two public functions with the same meaning but slightly different circumstances, having a protected method that handles both via an extra argument makes sense.
Consequently, the only way to end up with five conditional arguments on that function is to have an enormous number of public functions all doing more or less the same thing. Somewhere around 2, alarm bells should have been going off in your head, asking why you/we have done this and perhaps now is a good time to stop?
So at first you really have to dig into what “reality” is .e.g if you are building an auth system, how many ways do you login? Login page for customer, login via sso, login via google auth federated etc.
Now foreseeable reality is what is most likely going to happen. YAGNI proponents would say don’t do this but I’ve found that thinking through future scenarios that competitors are doing or customers asking for (even though you don’t have time to do it) should be taken into account. You realize it would be nice to support auth via personal tokens restricted to certain api calls.
Having a solid understanding about foreseeable future is paramount. Task #1
Then it’s understanding what pieces need to be stitched together to make it happen. What changes together, what changes independently. What are the separation of concerns between the pieces?
You need to represent different auth types in db. A central place that takes a request and it’s headers to authenticate. Also authentication and authorization are different things and shouldn’t be bolted in the same place.
Having an in-depth understanding of separation of concerns and boundary contracts is Task #2.
Nailing 1 and 2 down usually leads fo clean implementations.
Also huge believer of continuous refactoring. Sometimes reality changes. Focus on a small thing you understand for sure. Incrementally add things and re-think the separation of concerns and how reality is represented in datastructures. If that needs a change, make it happen. Building on wrong level of abstraction and working around it has caused a great level of pain.
One of the things that goes into it is how clear the direction the business is headed in and how experienced you are with both the tech stack and solving the specific problem.
Therefore, the opposite of the DRY issue could (should!) be Worth A Preemptive Unhooking of State. And most of that stuff will not be at all worth unhooking whatever state you've got in mind, in order to go into whatever layer of abstraction is required to DRY up something absolutely trivial and meaningless. Semantics matter. We need to be focussing on what's important in the problem set, not getting derailed into computational linquistics.
Code that is WAPUS-Y might be deeply unfamiliar territory to our more rigorous brethren, because the very idea of asking whether something is Worth A Preemptive Unhooking of State implies there are matters outside what we learn in comp sci, a subject to our coding outside the coding itself.
But that is so often true.
And maybe this framing of it can be, erm, sticky and illustrative in its own right :D
It’s a construct available for a programmer to use. DRY encourages one to do so when appropriate. Dogmas interpreted that one must always do so are simply unhelpful.
Even if the code is the same today, is there a case to be made for a new feature that would require them to evolve independently? If so, the "logical" meaning of your function is independent between the two use cases. The code looks the same, incidentally (by accident). Speaking moreso to "domain"/business logic.
And of course, there's the more obvious DRY case of truly context agnostic operations, like math functions, string manipulations etc. Ideally you're using some standard library for these things... agree with the author that it's better to avoid sharing tiny custom libraries if amount of logic is too low to warrant the extra abstraction. The cost of abstraction is not fixed, of course, but dependent on how common and familiar that abstraction is to the broader audience.
Even in the case of log2, some log2s might need to be fast approximations, others accurate down to the LSB, some operating on complex numbers, and others on array-like data...
Of course mathematics is always an endless game of expanding and contracting definitions so might need some additional versions anyway (most commonly something like f(x) = log(1+x)), though if you're clever about it you can rephrase most of what you want with existing functions.
So I just leave the mess for a while and see what will grow.
It is a perfectly good solution if we accept that refactoring is also a thing. And it should be. Code can not be perfect from day one.
These rules are can be applied widely. For example sometimes it is good to repeat code in s React component and leave it that way.
Especially if it is unknown how the design will evolve. So also it is sometimes better to leave some mess that is easier to manage than to slice everything that may make it harder to manage the code further on.
Often chasing DRY and other rules too early leads to trouble. I call it early optimization. Just let the code breath! Give it some time to make it perfect!
I wrote an article on this that came to similar conclusions: https://coderefinery.nz/2019/01/28/beyond-dry-why-redundancy...
Who says the cognitive burden is lowered just because of adjacency? If you have to keep N variables in mind, its about the same if they're at current depth or 1 level deep.
To me the most evil form of DRY is when you get functions calling functions with arguments that are just passthrough. If a function only uses an argument to call a subfunction, it's a sign that those functions should be decoupled and better orchestrated.
The hard thing to remember about code is, nobody gives a shit about your code when everything is working. Nobody is poring over it basking in your magnificence.
If they're looking at your code, it's because something is wrong. So they're looking for wrong things, and they don't know where the wrong thing is. So they're looking at every code smell they see along the way doing a probability calculation. Your ornate clever code causes them to dump important state, such as "why am I looking at this code in the first place?".
In the case of argument passing, we're all familiar with argument transposition. I have to check you haven't done it. If you rename variables, I have to watch for that too. If you have a stack of empty methods just ferrying arguments around, you've multiplied the surface area I have to think about.
So true until the executives are all like "Your team is not performing / low velocity" we're moving your job/team to $CHEAP_COUNTRY
Usually it manifests in having tens or hundreds of tiny, stateless functions, with very little logic in each one. There's a common line of thinking about this being "good" because you can "mathematically" prove the correctness of a function, and thus can ignore its implementation. But abstraction is only a big win when the abstraction is easily and widely understood. When you create a one off, ad hoc abstraction its purpose is not going to be obvious to the reader off the bat.
The important thing is that the total "api surface area" within your context is low, and each function/abstraction has sufficient meaning. There's a sweet spot where you can fit most of the local functions in your head, without having to constantly re-read to remember the purpose of a given function. Naming is quite important to this end.
The core conclusion I've come to is to take ALL such "rules" with a grain of salt. Programming is a skilled craft, and there are no rules that you can apply universally.
And then you can make the judgment whether the abstraction is an improvement over the duplication.
Better to have multiple lots of dumb code than one lot of clever code.
https://stackify.com/premature-optimization-evil/
Don't be evil.