I find considering whether the repeated code is innately related, or in the same domain, or not, to be the best metric for determining whether I should refactor it to not repeat. It's adhering to the principle of least astonishment. Sometimes it should all stay separate. Sometimes a few pieces should be refactored to share code, and others left to repeat. Sometimes one set of the repeated code should share a function, and then the other set should share a different function, -even if that different function has the same implementation-, simply because the second set is so unrelated to the first. Sometimes they all should share the same code, because they're all related.
Breaking things into small, reusable pieces helps avoid this because the dependencies you create are on small, easily understood, easily replaced snippets, which tends to have a very clear context/domain. It's easy to prove that the abstracted function is correct, since it's so simple, and it's clear what domain(s) it involves, since again, so simple, and thus the problem with any given bit of functionality is almost assured to be the composition of those functions, i.e., your specific use of it in this one instance. You can still run into issues if you reuse some of those composed functions, however.
In general, the greater the complexity and domain specificity of the code you're trying not to repeat, the greater the danger you're shooting yourself in the foot.
FP is especially beneficial, as such, because the higher order functions that tend to be common are both pretty simple in their function, and pretty universal in their domain.
So for example if you have something that is dealing with registering your website, create a directory called register_website and all models, functions and controllers to do with that code go in there. The talk suggested deleting the controllers and models directories. He also said there was duplication but everything was more loosely coupled and you could always group things in deeper functional units later.
I'm not 100% brave enough to go for this structure quite yet but it certainly makes coming to a project pretty clear as you literally have a list of all the functions of the site where the code sits for each thing :-)
In fact, my experience has been that for good design much more important than DRY code is simply reducing global variables (including classes -- which implies that you will be doing a lot of dependency injection). But, again, you can take it too far.
I think my biggest piece of advice for beginners is to get a mentor and constantly ask for opinions about whether a design is good or not (and what about it is good/bad). It's not something you can learn from rules. It requires a considerable amount of experience.
Loose coupling (which implies getting rid of global state) is at least equally important, as is failing fast.
I was reviewing some code yesterday and the developer had a bunch of interfaces for everything so she could swap out implementations and she was very proud of how modular everything was. The problem was that all of her interfaces accepted a File object instead of a stream so we could change what the behavior is but not what it acts on which is really the opposite of what we needed.
I find that most developers who think that they can predict what kind of coupling they will need for future requirements tend to be wrong, too, no matter how experienced. Only backwards-facing refactoring really works at producing a good result, coupling-wise.
DRY is generally meant to be an optimisation towards reducing the cost of change by reducing the amount of duplicate code (repetitive, busy-body work) but it often backfires and increases the cost of change.
For example, say you are tasked with creating 15 widgets. During your first few sprints, every widget is the same, so your colleague berates you for creating them separately. You refactor them so that they all share a common implementation and continue.
3 weeks passes and features are being added left, right and centre. A point of contention within your team is whether to differentiate between 3 different types of widget with different features, or whether to use dependency injection. Development is slow because only one person can work on the widget code at a time, and the widget code has becoming increasingly complex as the number of callers invoking it have given it increasingly diverse requirements (note how looking at a function generally does not tell you which other functions call it, how they call it or why they call it).
The question you should ask yourself in these situations is, how similar will these widgets be in 3 months time? If it's likely that they will diverge significantly then perhaps the best solution would be for them to be completely differentiated now: tested separately, easily updatable/extendable without regression bugs on other widgets, able to be worked on by multiple engineers without merge conflicts, etc. The unfortunate reality is that once you have made a unit of code DRY it is both time consuming and difficult to do this and politically untenable so people typically avoid it.
When should you make your code DRY then? (1) You should make your interfaces DRY early on as this is a safe win without the problems of DRY implementations, (2) When you are very sure that the implementation of different units of code will not change in ways that do not create intellectual work of the following kind: "How do I avoid complecting these unrelated behaviours?"
Having said all of this, perhaps this is not beginner friendly. Perhaps beginners should just follow DRY always. At the very least they will get practice refactoring.
DRY should only be avoided if the refactored code ends up being much more complex than the duplicated code and it doesn't reduce the duplication all that much.
Pretty much every time I've seen DRY used to eliminate repetitions of three or more times it's been a good idea and 4 or more times it's been a no brainer.
>The question you should ask yourself in these situations is, how similar will these widgets be in 3 months time?
This is an awful approach. You won't be able to predict how similar they will be and you will likely create expensive to maintain architecture for projected scenarios which don't exist and neglect the architecture for scenarios that will.
I've seen many codebases littered with architectural changes that were put in place for presumptive future features and then never used. It happens way, way more often than those architectural changes actually being used.
> > The question you should ask yourself in these situations is,
> > how similar will these widgets be in 3 months time?
>
> This is an awful approach. You won't be able to predict how
> similar they will be and you will likely create expensive to
> maintain architecture for projected scenarios which don't
> exist and neglect the architecture for scenarios that will.
I'm not talking about writing code for features that don't exist. I'm talking about not doing extra work to generalise your code too early. DRY is generalisation. It is in a sense 'writing code for a projected scenario in which a feature is the same for all of its callees'.Keeping things specialised towards a single purpose even if it is redundant in the short-term makes code easier to grok, test, and update.
Making a change across 5 redundant functions is not time consuming work and is not easy to mess up. However making a change to a singular function which needs differentiated behaviour for one of its callees is error-prone as (1) it might force an interface change and (2) there is a non-zero probability of a regression error on the other four callees.
I'm not attempting to predict the future. I'm attempting to talk from first principles about real cognitive and communicative costs.
I started using a rule of "use the language's primitives first" to guide my use of DRY. There are many ways in which you can "reduce repetition" by introducing a generalization abstraction, something that in turn couples lots of code to a master definition: Generics, high arity functions, and inheritance are three such temptations. Doing that, in turn, accidentally models the whole problem before the spec is mature.
But if your discipline is "I'll use more built-in types and simple data structures, I'll keep more of the code inline, I'll let this class get bigger and I'll make some duplicated simple, fixed-function code paths," you don't get trapped by that. Sometimes I end up with the same simple math function duplicated across two different modules because I needed it for completely unrelated algorithms: That isn't pretty, but it's not a gross failure.
The opportunities to make a really useful abstraction for a novel problem often take time to ripen, and until that point you might as well assume that the code necessarily involves a "spaghetti" mindset. You don't want to pay the cost of reworking the abstractions unless you have to - it presents a huge hazard to maintenance and it means the attempt to generify failed.
But I think coders do get caught in a "clean code" mindset now where it needs to look both pretty and succinct, 100% of the time, and that leads to pushing the functionality out of the function body and into "ravioli" code.
I think Go is a good language, of the hot ones today, for learning this kind of engineering-centric focus because it actively discourages trying to whittle down the code to the succinct one-liner. Unfortunately by the time I encountered it I already had this figured out(perhaps years later than I should have).
e.g.
widgets/
simple/
index
test
advanced/
index
test
weird-as-hell/
index
test
company-x-wanted-different-features/
index
test
tophats/
index
test
In the structure above, each of the widgets are tested separately, each are very easy to find, and each might have redundant code doing the same thing.This is way simpler and easier to work with than if people have created an `uber-widget` that has a configuration object and some dependency injection to handle all of its features because some overzealous engineer once decided that he was certain that all of the widgets would be able to be represented by a fully generalised `Widget`.
I'm obviously not saying that all cases of DRY are like this, but I am saying that there are far too many people that end up doing this because they make decisions to generalise code too early.
Bad repetition is when you or other engineers have to also remember to update the repeated parts.