Why do so many developers get DRY wrong?
changelog.com
changelog.com
This type of topic is hard to talk about. It's so nuanced that saying a statement about how to do it sounds like a a gutless generality. It also depends on the programming language and lifetime of the project. I've banged out some real ugly code when servers were on fire, but it was all stuff that was destined for an early death.
Both internal and external APIs must be kept coherent.
Also when readers are trying to understand exactly how some function changes the system state, having to refer to numerous other functions it calls is tedious.
That's a really good way of putting it.
I try to avoid it by not writing new functions, but gladly using existing ones, in my implementation of whatever single new one.
For example if I'm writing a find_and_update_foobar function, I'll use find_foobar if it exists, and with the right signature, but I won't write it just to implement the one I actually care about; ditto update_foobar.
But, professionally I've mainly only used python; so I find it still deteriorates into a mess. (I type hint extensively, but still it only takes some missing hints, or something too loosely - or wrongly - typed.)
I haven't used rust professionally/enough/on something large enough to be sure, but my feeling is that it just having a type checker prevents so much mis-refactoring.
Are there things you may only do iff you using a single function body?
Perhaps it is not always a bad idea to decompose that function into 2 properly named "steps", which reside in their own functions.
This also applies for small inlined common helper functions. If it folds two or three operations into one, but doesn't have an obvious universally recognizable name, it would often be way clearer to a reader of the code to just put the few operations in directly.
Yep. That's why changing system state is better kept at a minimum, and at the highest layer possible.
Most languages have a very clear public/private marker for determining what is an API and what is just there for convenience (it doesn't have to be literally public/private, for some it's scope, others just tag the names, etc). You don't need to change a code's style just because of it.
On the other hand, if the functions do not have logical and atomic meanings, they will do make your code as hard to understand as it will be to name them.
Not when you work on a team that does this as a practice.
Many teams don't manage to have a common code style, which is a big problem in itself.
> Also when readers are trying to understand exactly how some function changes the system state, having to refer to numerous other functions it calls is tedious.
The point of breaking out the small functions is to name them so understanding the code becomes easier. This takes some skill and thought, of course, and if done thoughtlessly, it will not be good. But that's true of every practice...
Overly granular functions constructed to name blocks of code are a firm anti-pattern in my book.
I'm sorry you've been in these teams. There are good teams out there!
The oldest codebases I've worked on were 30 years old; the one at my current company is about 10 years old. People come and go, refactorings are started and stopped, bugs are fixed on a tight schedule, and the structure of code divided into nominal blocks slowly dissolves.
Coalescing state transitions should trump decomposition. But it’s often the case that you can do both at once.
Don’t make me look under every rock to figure out why you changed the stuff I gave you.
We aren’t built to deal with chaos. When you put stuff in one end and something random comes out the other end, you have no idea at all what’s going on (until you do, and then they try to put you in charge).
You want heaps of code that looks but doesn’t touch. Push the changes to the edges where they are easy to see. If I can trust that things in the middle aren’t mucking around with things all the time then lots of little methods don’t hurt my ability to think about what is happening.
(This is not quite the point of Hexagonal Architecture, but they have some common ground. See also early Angular’s philosophy of cooking all input data immediately and passing it around cooked.)
The more of your app you can write as pure functions, the more of it is predictable and easily testable, so you don't have to worry about how it's going to behave.
That's part of why we now specifically recommend writing as much of your logic as possible in reducers:
https://redux.js.org/style-guide/style-guide#put-as-much-log...
In a purely pigeonhole principle sense, more pure code crowds out impure. But if you can put all of the impure code into one or two phases of an interaction, you’re improving things without really changing the ratios at all.
We often get sucked into analogs of the things we really need and I feel like this is one of them. You want the state changes to be comprehended. Fewer of them might help a lot or make things worse (legibility or performance). I believe it’s another quality over quantity situation.
This actually helps to speed up navigating code, because parts of it get meaningful names.
This is the common justification, and it's misguided. It turns out that you need to be aware of the "details" every time you look at the code.
If, by looking at the code enough, you memorize the "details," it's tempting to move them to a new function with a clever name that tickles your memory. That won't help anyone else.
Use comments to introduce blocks of code that need explaining. Use functions in coherent APIs.
No, I don't. I need to be aware of the relevant details, but code in functions with meaningful names mean (1) I can skim a high level overview to see where the details relative to my current interest we likely to be faster and, (2).I can zoom in without distraction to those more easily.
> Use comments to introduce blocks of code that need explaining.
Comments make a wall of code that is already hard to get an overview of because of its size less legible, decomposition did the opposite.
Yes, and decomposition provides an outline, which is faster to skim than an article. There's a reason tables of contents are a thing; they let you find wheat you care about much faster than linear text with section headers.
With code, they also have the advantage of allowing you to make use of step-over/step-into debugging.
That's less the case if the functions are local (either not exported from the module they are included in or actually function-local, in languages supporting nested functions, to their caller.
> Both internal and external APIs must be kept coherent.
If it's not external, it's not an application programming interface. And, even ignoring that, there is no definition of coherent for which that approximates truth that is inconsistent with single use functions for code clarity and organization.
Even languages that don't have mechanisms to make methods/functions truly local/private tend to have conventions for indicating functions/methods not part of the intended-as-public API of a class or module.
> Also when readers are trying to understand exactly how some function changes the system state, having to refer to numerous other functions it calls is tedious.
Conversely, I find the code having functions with descriptive names helps me not have to read through a bunch of irrelevant code when I'm doing that, which helps me get to the part I need to understand whatever I'm trying to understand faster and with less distraction. Yes, if it's done badly it's problematic, but that's true of literally everything.
EDIT:
> Also when readers are trying to understand exactly how some function changes the system state
Yes, if you are doing relatively unconstrained imperative code that's willy-nilly modifying state, breaking that up into functions passing mutable state around is likely to make it even more incomprehensible. Decomposing into units that are externally-pure functions (that is, while they may do local mutations, they don't modify any state received from the caller), though, is not problematic. In general, except for subsystems dedicated to managing mutable state (and where this function is kept as constrained as possible), I prefer coding in externally-pure functions in general. When you start with that and use it as a constraint on decomposition, decomposition can no longer serve to obscure state modifications. Indeed, it clarifies and more narrowly isolates them.
Splitting hairs, I think. Even if not intended, someone else may assume your "documentation" function was meant to be generally accessible (i.e. an API), and call it that way.
And then, suddenly, instead of a single use function that serves mainly for documentation and code organization, it js a locally resuable and locally reused piece of functionality. Which is a problem, ...why, exactly?
The only difference between a well-designed single-user function and a well-designed locally reused function is that the single, clear, and unit tested purpose of the single-use function isn't needed in more than one place when it is written. If that changes, why shouldn't it be reused?
Refactor with a new internal and coherent API once it's clear that's needed.
1. Use them if they take <5 parametets - no point making a function if you need to pass all the local context into it
2. Use them iff they return something that can be clearly named and have an explainable and understandable TYPE, even in dynamic langs where you don't explictly type things - if you'd end up returning a tuple of 3 arbitrary things that can't be "named as one single concept" it's a code smell
3. Kepp those functions local, maybe define them inside the calling function, at least private if they're methods, anything - if your single caller function becomes unwillingly part of an api, which can esily happen in Python with a _funk() method that users could ignore it's private, you've lost
procedure foobar;
procedure helper1;
begin
end;
procedure helper2;
procedure subhelper2;
begin
end;
begin
end;
procedure helper3;
begin
end;
begin
end;Only the foobar function can access the helper functions.
In order to do this, recognize iterations are key and that code must both work today and allow for future evolution.
With clearer goals, one might save time and effort refraining from premature optimization and overzealous refactoring. When direction is unclear, you can comfortably lean on established accomplishments while avoiding regressions.
DRY and YAGNI may thus take a backseat to incrementally deductive design.
Also, I think we kind of said goodbye to DRY after we decided that microservices is The Way. How many microservices reinvent the same goddamn wheel?
Trying to nail down DRY abstractions not too early but not too late is often tricky but concern over it probably puts you in good company.
1. Write code 2. Cleanup the mess 3. Write more code 4. Realise there a better way. 5. Stick with it because effort to change < benefit gained.
In Refactoring Kent Beck doesn’t just show extracting variables, functions and classes but also re-inlining them.
The problem is at scale, re-inlining is hard. If it were easier to do, it might be done more.
I find that waiting for N repetitions of some abstractable code is a heuristic to mediate between the cost and benefit of refactoring.
I don’t have an answer for improving it. Perhaps making easier to automatically manipulate the code in this manner?
I’ve thought for a while that it would be interesting to treat a code base as a database. Applying documented migrations to it which describe the code transformations occurring (and how to reverse them). Then changing these aspects of code would be trivial.
But then your meta program, the set of transformations which generate your current code base, would probably suffer the same problems as your codebase.
let (|>) x f = f x
It’s extremely simple, extremely powerful and it’s used pretty much everywhere.DRY falls into a different category. For example if you use `.toFixed(2)` to format floats everywhere it is better to abstract it to a function. And I believe this is what the article is about.
Formatting a float is a kind of logic or knowledge about how you present values in the view. You should not repeat this knowledge because if it is later decided that the view must show 3 decimals you are in trouble.
Maybe DRY can be easily explained by: don't repeat (business) logic.
It's the source of a large portion of the accidental complexity I find in code. "If I just create this abstraction, all this duplicated code goes away" - we've all heard it and many of us have told it, but few of us realise that it's the prequel to the most popular story of all: "all this code is such a mess, there are all these extra layers that don't really make sense and unpicking it is such a pain, I can't believe someone wrote this".
The story inbetween is about a young, inexperienced developer who has 3-days to deliver the one-feature-to-rule-them-all, to appease the almighty project manager, necessitating an adventure into the labyrinth carefully crafted by the developer in the first story.
I'm always amazed at how eager people are to over-engineer a solution that makes it a mess to deal with moving forward. Developers at large like to appear clever, tend to have (fragile) large egos, and don't seem to want to veer from established dogma--much of which based on little evidence or evidence that doesnt apply to a case they're dealing with.
For me it's usually, "Oh crap that thing I changed I had to change here and here too, whoops its good now... Wait no I also had to change it here... and here... now that we're done with that we should be fine... DAMMIT!"
It looks like it is, you'll see plenty of people claiming it is, it may even lead people into solving the problem after they gather some experience. But it's not what solves the problem.
- keep data serializable
- cluster transformations accordingly (Keep things that belong to each other as close as possible to each other, in a literal meaning, same file/module/etc. A new coworker should have the feeling of entering a public library - she'll know where to look for.)
- it will probably never be the case that you have to invent a new data structure or algorithm
Developers understand such terms like 'function' in very different ways (see top comment in this thread). Some have a more abstract approach and understand it as 'pure', where others have a more instruction-bundeling approach and understand it as 'procedure'. Either is valid, but are completely orthogonal to each other when it comes to its usage. I understand Gary Bernhards approach to 'functional core, imperative shell'[1] as a mean to talk about this - instructions demand for composable building blocks, and those building blocks come either from:
- built-in functions (with an already universally understood API)
- or your own (clustered) utils (and demand an easily understandable API).
[0]: https://mitpress.mit.edu/sites/default/files/sicp/index.html
[1]: https://www.destroyallsoftware.com/screencasts/catalog/funct...
parseDataFromHTMLResource :: Resource -> HTMLRaw -> Maybe Data
...and assuming we validate the HTML: validateHTML :: HTMLRaw -> Maybe HTML
I think your question relies on the separation of 'Resource', as a domain separation like right below will probably create a lot of duplication (since any resource is free to structure their HTML within the standard however they want): parseDataFromGoogleHTML :: HTML -> Maybe Data
parseDataFromGithubHTML :: HTML -> Maybe Data
parseDataFrom...
Perhaps you can reduce the duplication by converging to different fingerprinted resources. So a HTML resource, fingerprinted by style, will then guarantee your data: data HTMLFingerprinted =
HTMLStyleGoogle
| HTMLStyleGithub
| ...
htmlFingerprintedFromHTML :: HTML -> Maybe HTMLFingerprinted -- So the Maybe will be here
parseDataFromHTMLFingerprinted :: HTMLFingerprinted -> Data -- ...not here
(Note that this may be what the parent means. It helps solving "For me it's usually, "Oh crap that thing I changed I had to change here and here too, whoops its good now... Wait no I also had to change it here... and here... now that we're done with that we should be fine... DAMMIT!"" with "I just create this abstraction".)I would say if you're able to implement the function 'htmlFingerprintedFromHTML', you deserve the abstraction and thus DRY the codebase on the fly. And this "no pain no gain" mantra is what I personally really like on a 'functional core'. Code stays only duplicated where absolutely needed until you find a solution for the abstraction.
...and I guess implementations of fingerprinting HTML may vary a lot.
DRY is a good learning tool, and it is an after the fact property of good code. But people shouldn't ever preach it.
The dominant example of this in my mind comes from some traumatic (and dramatic) work experiences involving web scrapers.
You scrape pages A and B. Both require logins. You notice that A and B use similar code to login, so you factor it out and now you have "connected" A and B. Or, more accurately, the common "login" method is connected to A and B.
The problem is that A and B are separate ships, and they are going to different places, and you tied a rope between them. That rope is going stretch and fray and break. There's no reason logging in to A and B should be similar, it just happened to be that way at one point it time. One day, probably, those two sites change their login workflows to be completely different.
So "just repeat yourself" and let each scraper be self contained. Let each ship sail their own way, don't connect them.
One case of this isn't bad, but if your not careful you end up creating dozens of connections between things going all different directions. Break your project into "things" and graph their dependencies. If you can break X by changing Y, then X depends on Y. If you can break Y by changing X, then Y depends on X. Your dependency graph should look more like a tree than a total graph.
Unreasonable X is unreasonable, that's true no matter X but that's not very insightful. Unreasonable factorizing is unreasonable. So are unreasonable copy pasts.
Now a good question can be: would I "prefer" one or the other gone too far? That's merely a personal preference question, and on my side I've got a clear answer: I prefer something a little bit too abstract, to something a little bit too copy pasted. Because while abstract I still can understand more of the system more quickly, whereas duplicated I basically have to start by reading a lot more and factorizing (possibly in my head) before getting to a balanced situation...
Your own preference may vary, but I've yet to see a system maintained correctly by people copy pasting too much and not caring about anything but the one example they have to handle (because e.g. of a workflow described in a ticket).
Things can be accidentally similar, but WAY more often in copy pasted code-bases they are accidentally diverging, and that multiply the time I spend on that kind of mess by a factor that is probably near 10.
[0] where soon either means really soon because product management needs it yesterday or never because there is too much work to do in other places.
but in a way that just pushes the problem to coupled interface signatures..
now I wonder, is there research in variability based interface design ? looking at systems and estimating how much some parts can change and how much are probably easy to pin down forever ?
If you construct your program via the point free style using function composition, the dependent function can easily be swapped out of the composition and replaced with the modified function that you need while maintaining DRY to the maximum possible effect.
You are talking about a fundamental question of program design that is largely solved by functional programming via the point free style.
It is talked about here:
https://www.quora.com/Is-senior-full-stack-engineer-Ilya-Suz...
This would be a variable that can be muted. More commonly, when people refer to "mutable" variables, they mean variables that can be mutated. It's not clear what muting a variable would mean.
From context, hopefully other humans will recognize the slight error and comment on the topic rather than walk away totally befuddled and comment only on the error.
Do you not understand what I am saying in the comment? It's too late for me to edit the comment, but I hope the meaning is clear despite the slight error.
"Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize."
Around 30 min, Rich Hickey describes how the opposite of simple is complex, and mentions the etymology, "to braid together".
"It's bad. Don't do it."
Or, you use the same code for all and if the sites change the login flow then you add a custom one for the site which is different.
Fundamentally, I don't agree with the argument that I should copy and paste code now to avoid the possibility of copying and pasting the same code later.
The base case of retrying is `zero retries` (or `one try`).
In each of the 10 scrapers?
If it's common code for the retry, now all my login code needs to raise the same type of exception or otherwise signal in the same way that the login has failed. Except instead of the abstracted retry code needing to work with one login function it needs to work with ten independently maintained functions that happen to be the same currently.
Or, in my opinion, handling it in the typical way functional programming does, you'd have your stateful computations like login represented as functions returning IO, which you can easily use off the shelf functionality to rate limit and retry (like cats effect, fs2, etc... In scala). This kind of programming isn't as mainstream as it could be, but if you can build retry once and use it for pretty much any side effecting computation, you wouldn't feel a need to DRY up things that should be separate in an attempt to share code.
I think this article does have something going for it. DRY should be about knowledge. Don't repeat yourself by handling tax rates all over, get that into a central place. This is not about turning similar looking blocks of code into a clean single block that handles everything as these tend to actually hurt maintainability (ever see the littering of conditionals in "DRY" code because multiple call sites use the similar code slightly different? Yeah, you did it wrong).
Another comment wrote it well by mentioning SPOT: Single Point of Truth. DRY and WET SPOTs. I feel like an analogy is forming that someone more quippy than myself can ferret out.
This is where DAMP comes in. Descriptive And Meaningful Prose.
I should be able to read the test failure message and know what I broke. Barring that, I should be able to read the test (not the whole fucking test file) and tell what I screwed up.
Anything else kills momentum, and I just want to get away from your code as fast as possible. Which means more tech debt.
If you google for it, you will find the synonymous “Single Source Of Truth”, which however makes for a worse acronym.
And so "DRY", to the extent that it's useful, encourages you to find slack areas in the code where there's low potential for introducing coupling, and to factor those out so that you have code that is mostly-decoupled without also being redundant and hard to modify - the factoring reflects "knowledge" about the problem. And yet it's not always obvious when you have the knowledge or not. Sometimes redundant-looking code is a form of hardcoded data and a factoring would only push it towards being fully data-driven(which exacts a price in debugging). The Rule of Three is just a common way of making this decision about knowledge.
Also what have globals anything to do with that, and why do you put them in the supposedly factorized code. You are mistaking it with a bad mess you once saw, maybe? On the other hand, code bases obtained through copy paste based programming can not be considered anything else than a bad mess.
But yes, factorizing can be done badly, even to the point of being counterproductive. Like anything.
I think they're assuming that everyone has read The Pragmatic Programmer. To quote the original DRY principle from there:
Every piece of knowledge must have a single, unambiguous, authoritative representation within the system
Note that this has only an accidental relationship with code duplication, and in some cases could increase the latter.
It often is an implicit business rule that links several statements together or a system requirement like freeing db connections and locks in the right order after use.
If there is only one right way, there is no reason for change or changing together, so is not shared knowledge.
The most important thing isn't to apply these heuristics, but to understand the problem space in which your code operates before you lay down your abstractions. That's a difficult thing to do without domain knowledge, and in a lot of enterprises you will never get access to the kind of domain knowledge you need to refactor effectively unless you're the lead or in management.
To the extent that your code is a series of statements about how the system behaves in response to a particular data input it's easy to read and documents itself. And to the extent that your data structures and statements resemble statements that a domain expert might make(move gantry 30 meters to the left, then drop the crane) they become easy to change in response to changing requirements. Domain knowledge tells you what the fixed elements of the problem space are(is it always a gantry? Does the gantry move in any directions other than left? What does it mean to drop the crane and do we do it different ways?). That informs how you structure your code and what the most clear factoring is.
It will not always be the smallest refactoring.
In designing a web page, if you find yourself saying "there is a button here" in HTML, and in CSS, and in JS, and on the back end, that is not DRY even though the syntax looks nothing alike.
Two different APIs, for services serving different purposes controlled by different external entities, happen to have the same structure and you find that a large chunk of code can be factored out of both. I would argue that "what we currently need to do to API 1" is a separate piece of knowledge from "what ... API 2", and unifying them is not DRY.
If I encounter it a third time, then I've got enough data points to make a good guess about what the right abstraction will be. If I've done a good job so far, it shouldn't be too difficult to refactor it. (Strong, static typing helps.)
This is, of course, just a heuristic, and it's not all-or-nothing. I'll take my best guess about what the right abstraction is going to be, and I'll try to get it right the first time. The second round also presents opportunities to take two points and extrapolate a line.
It all comes down to experience: not just with the system, but with the domain that the system is about, and with the way systems change and grow. No one rule of thumb ever encapsulates all that.
The third time I automate it. By then, I understand it well enough to have good odds on being able to do the automation successfully.
If it's more than three times, you ought to automate the automation!
Part of the point of doing it the second time is to make sure that I really understand what I'm doing and how I'm doing it. Without that, I can't write the automation on the third time.
Well, if the task is "automating things" (for very general values of "things"), I don't understand how I'm doing well enough to automate that.
Premature abstractions are way worse than repetition. A poor or insufficient abstraction leads to obfuscation which leads to misunderstanding which leads to novel constructs for the same responsibility. Because a poor abstraction can be really really difficult to back track, you end up with hacky work-arounds to get something done.
I think encountering novelty in a codebase is the biggest thing that damages comprehension; and repetition actually enhances comprehensibility.
> The trouble with DRY is it has no reference to the knowledge bit, which is arguably the most important part!
Okay. Now what does this mean? Is this article effectively a tantalizing recommendation to read The Pragmatic Programmer?
More that it's a guideline, not a law. We should always use best judgement to decide when the tradeoff of readability and declarative code is worth a small amount of repetition, rather than religiously refactoring something for the sake of it.
It's a bit difficult to get across in text, but the minimum number of repetitions of a piece of code to make it "worth" putting it in a function is... 1. (According to me, and Tony van Eerd of Postmodern C++ fame. I had come to this conclusion on my own, but his talk really articulated it well.)
It's all about limiting the scope of side-effects, accidental reuse or variables, etc. etc. such that a human can do chunking to understand the whole.
I generally find that this is not an easy thing to capture in "metrics" or "rules". Guidelines with reasonable rationales, etc. etc. and when-not-to's, definitely, but that's a really hard thing to do and it doesn't get many clicks.
EDIT: ... and just to get back to DRY. The acronym is far too absolutist, but Try-Not-To-Repeat-Yourself-Too-Much-Unless-You-Have-Good-Reason-To isn't quite as catchy, is it?
One concrete example: If your software has to create really complex objects, would you rather describe _how_ to create those objects in 10 places or one place? That's a scenario where you don't want to repeat yourself.
Dan Abramov [wrote about](https://overreacted.io/goodbye-clean-code/) this (linked in the OP), but in his example he's removing repetitive code. He's not removing multiple copies of the _knowledge_ about what the program is supposed to do.
It's a subtle difference that seems more difficult to describe than I'd like, but it's an important one.
So for example, documentation (truths about the code) should derive from the code (eg by doc generation). Otherwise the docs & code will drift apart. Or if you're passing domain information across the wire between client & server, you should derive the data structures at both ends from a common source.
I don't get it. Code /IS/ knowledge and whenever I copy-paste code around, I duplicate not only code but also knowledge.
So you have an API that belongs to this project. When you change it, do so in one place, and then run your doc generator rather than change it in both function/method signatures and docs.
Or you have domain knowledge embedded in classes, and a wire protocol between peers or client & servers using different languages. Derive the data structures in the two different languages from common source (either one of the languages, or both from metadata).
I think the distinction between this kind of project-specific 'knowledge' and more abstract 'everything is knowledge' issues is clear enough in practice. But it is just a rule of thumb rather than a deep philosophical principle, and like all such will break down in individual cases.
whenever I copy-paste code around
But that's just one source of code duplication. Another might be (for example) duplicated code deriving from code generation. DRY might advocate this (as there's a clear canonical source of knowledge), whereas a generic rule against all 'duplication' wouldn't.
In that case, there is not a single piece of knowledge being duplicated, but rather two separate pieces of knowledge being possibly unified.
Like, if you saw Http.getClient(...).doGetRequest(...) a few times, it wouldn't be worth pulling them out into a myGetRequest(...) method. Your teammates already understand the existing, repeated statements, but they haven't seen myGetRequest(...) before, so you wouldn't be making the code any more readable to them.
But if you had Http.getClient("auth.myservice.com:8443").doGetRequest(...) in a few places, then I would pull out the host and port (or maybe the whole line), since it contains knowledge of where/how to authenticate.
Coming from the other direction: if I were reading the code, I can imagine myself looking for the one place where the auth happens, but I can't imagine myself needing to know the one place where Get requests happen (even if the 'Get' code is repeated much more than the 'auth' code)
I think the deeper problem in the software industry is that we have no collective memory and need to DRY up as a whole.
We should be weary of micro-optimizations for "elegance" which actually hurt the larger-scale maintainability of the system.
Though an alternative to (1) is that the meaning of DRY in common dev parlance has changed & has come to mean something different from Thomas & Hunt's intention.
DRY, WET, DAMP, SPOT, KISS, YAGNI...
Software engineering is too nuanced to be summed up in an acronym and surely the inventor of each acronym only intended it to be a basic rule of thumb.
let setX = (e) => {this.x = +e.target.value}
... other setup ...
input({onchange:setX, value:this.x})
This code is not a subject of undry or rule of three, as it’s ratio of meaning to character count is too low.And yet some frameworks make a decision to abstract it out at a wrong point:
let [x, setX] = ...
... other setup ...
input({
onchange:(e) => setX(+e.target.value),
value:x})
instead of clean and readable input_num(this, 'x', {})
// or
this.input_num('x', {})For example, if you reuse the same logic in a couple places, where the only difference is some specific variable, it should be written as a block with alias variables at the top. That way, two different cases look literally the same (except for a couple assignments at the top).
The question is whether it is a coincidence or the same concept.
"DRY" is catchy, easy to talk about, easily verbed as a recommendation, seems like a good idea, and seems to be recommended by people who know what they're doing.
The problem is that it seems self explanatory, so no one discusses the definition. At the same time, the more obvious definition isn't the right one.
As a counter-meme, I have been proposing we refer to overaggressive syntactic deduplication as "Huffman coding".
To determine what it is, you need to understand the domain. Part of that is knowing not just what it is, but how it generalizes and predicting how it is likely to change.
Of course, this is approaching perfection. One could also cut and paste monkey-like, e.g. instead of looping (without a performance need to unroll loops).
IE, if you wrote some code in one place that you needed elsewhere, copy pasting the code is fine under my interpretation because it allows you to spent the least amount of time working on a problem you've already solved.
what I've learnt throughout many years of coding is that purest mantra one should follow is YAGNI. devs think like devs and strive for perfect code. but that goes directly against the business. the majority of the entire world is being run on really bad code. but that bad code works. and that is what is important.
in a way, you should treat things like blackboxes with strict interfaces. no one should care how the box works, as long as its interface works like it is supposed to.
PS: the dangers of DRY are introduction of deep dependencies that might, and probably will, bite you in the ass along the way. DRY should be used only for libraries, not for business logic - ever.
As it is, it left me with the same thought as those who claim to never need debuggers or object-oriented features: Fine let's say you're right - how do I implement your system?
Code is closer to a craft like carpentry than a pure knowledge job. There aren't any rules, only heuristics, and it takes time to hone your skills (not "learn" them).
The biggest limiter in this is the excessive tendency for flat hierarchies in dev. It flies in the face of the apprentice - journeyman - master system that has always naturally structured the delivery and learning of craftsmanship.
Per my understanding, it's anything you might say about your system, especially as it relates to your domain.
"There is a button here", "this is what we store about users", "broken widgets are red", "this is how we calculate interest", ...
They always must be leavened with Good Judgement.
I once had the (mis)pleasure of dealing with some novice DBAs. There were instances in the data model where things were so recursively linked that one query would fan out to 200+ to actually return the set of meaningful data. As you might of imagined, this caused scaling issues when you started dealing with significant amount of data (hundreds of GB of data in the DB).
Eventually we brought in some pros and one of the first things they did was eliminate a number of overly recursive queries and duplicated some (not all) fields of data to speed up performance. We went from 1->200+ fanout in a typical query to 1->25-ish. The performance gains were insane.
Of course it violates the "don't repeat yourself" and to abstract linked (repetitive) data in a relational way. But sometimes this best practice can really be counter-productive in edge performance scenarios. But in general yeah, don't be repetitive, abstract your stuff, and keep it clean and easy to change globally if/when needed.