To me, it's the things that are specifically intended to behave the same should be kept DRY.
If you're copying and pasting something, there probably isn't a good reason for that. (The best common reason I can think of is "the language / framework demands so much boilerplate to reuse this little bit of code that it's a net loss" — which is still a bad feeling.)
If you rewrite something without noticing that you're doing so, something has definitely gone wrong.
If a client's requirements change to the point where you can't accommodate them in the nicely refactored function (or to the point where doing so would create an abomination) — then you can make the separate, similar looking version.
I would embrace copying and pasting for functionality that I want to be identical in two places right now, but I’m not sure ought to be identical in the future.
If two countries happen to calculate some tax in the same way at a particular time, I'm still going to keep those functions separate, because the rules are made by two different parliaments idependently of each other.
Referring to the same function would simply be an incorrect abstraction. It would suggest that one tax calculation should change whenever the other changes.
If, on the other hand, both countries were referring to a common international standard then I would use a shared function to mirror the reference/dependency that they decided to put into their respective laws.
Sure, we could take the Foo, Bar, and Baz tables that share 80-90% of common logic and have them inherit from a common, shared, abstract component. We've discussed it in the past. Maybe it's the better solution, maybe not. But it would mean that instead of maintaining 3 component files and 3 test file, which are very similar, and when we need to change something it is often a copy-paste job, instead we'd have to maintain 2 additional files for the shared component, and when that has to change, it would require more work as we then have to add more to the other 3 files.
Such setups can often cause a cascade of tests that need updated and PRs with dozens of files changed.
Also, there are many parts of our project where things could be done much better if we were making them from scratch. But, 6 years of changing requirements and new features and this is what we have - and at this point, I'm not sure that having a shared component would actually make things easier unless we rewrite a huge amount of the codebase, for which there is no business reason.
What made your team decide on that rule? Could your team decide to drop it since it hinders improving the design of your code?
Thanks, I'll have a think about it
Copy-paste is free; abstractions are expensive.
https://www.youtube.com/watch?v=8bZh5LMaSmE
Worth watching in its entirety, but the quote is from ~13:59 in that video.
It explains so much of what has been bothering me about what I work on at work, and now I understand why and some of what to do about it.
Otherwise I mostly agree.
As an analogy, when writing a book, it's the difference of not repeating the opening plot of the story multiple times vs replacing every instance of the with a new symbol.
There's no one rule. It takes experience and taste to make good guesses, and you'll often be wrong even so.
Likewise if someone notices a bug in the method, you then have to go through and figure out which copies have the same bug, and fix each one, and QA test each one separately.
The proper approach is to make a judgement call based on how naturally generic the method is, and whether or not the existing use cases require custom behaviours of it (now or in the near future).
Even the most obvious of functions like sin() and cos() may in some circumstances warrant a specialized implementation. Sure, for most stuff you should not have 10 copies of those all over the place. But sometimes you might.
DRY is a bad rule. The more appropriate rule is avoid duplicating code when not doing so results something better. I.e. judgement always trumps rules.
Somewhat off-topic, that's one usual failure mode of "DRY" code. Code is de-duplicated at a visual level rather than in terms of relevant semantics, so that changes which should only affect one path either affect both or are very complicated to reason about because of the unnecessary coupling.
This is unhelpful even if the design is a complete mess.
Now there’s a new requirement that only applies to Europe and nowhere else and it’s super easy and straight forward to change the infrastructure.
I don’t see how it was a poor choice to literally copy and paste configs that result in hundreds of thousands of lines of yaml and I have 25 yoe.
Downsides with your approach is:
1. Now whenever you want to change something both in Europe and (assuming) USA you have to do it in 2 places. If the change is the same for both, in my system, you could just update the default/shared config. If the change is different for both it's equally easy, but faster, since the overrides are smaller files.
2. It's not clear what the difference is between Europe and USA if there is 1 line different amongst thousands. If there are more differences in the future, it becomes increasingly difficult to tell the difference easily.
3. If in the future you also need to add Africa, you just compounded the problems of 1. and 2.
From an operational perspective it’s much more important to ensure the code is clear and readable during an incident.
Overrides are like inheritance. They are themselves complex and add unnecessary cognitive load.
Composition is better for the common pieces that never change across regions. Think of an import statement of a common package into both the Europe and North America folders.
I easily see the one line diff among hundreds of thousands using… diff.
Regarding Africa, we’ve established 1 is a feature and 2 is a non issue, so I’d copy it again.
This approach scales both as the team scales and as the infrastructure scales. Teammates can read and comprehend much more easily than hierarchies of overrides, and changes are naturally scoped to pieces of the whole.
By way of analogy for why the two configs are different, for example Two beaches are not the same because they both have very similar sand.
You really have two different configs.. You also have one set of configs. You didn't set up an application that also fetches some config that is already provided. It would be like having a test flag in both config and database, sane flag - two places.
Where config duplication goes bad is when repeatedly the same change is made across all N, local variations have to be reconciled each time and it is N sets of testing you need to do. Something like that in code is potentially more complex, more obviously a duplication of a module, just more likely to be a problem overall.
Ok, so make a build script that creates a complete copy for each separate config from the default config and the overrides. Then use the complete copies.
> I easily see the one line diff among hundreds of thousands using… diff.
Yes, and in 5 years when it's not just 1 line?
Your approach does NOT scale. You just haven't scaled yet and haven't realized this.
Perhaps one day you will. I'm a dev who worked with infra people who had your philosophy: many copy pasted config files.
Sometimes I needed to add an env var to a service. Expressing "default to false and only set it to true in these three environments" took changing about 30 files. I always made mistakes (usually of omission), and the infra people only ever caught them at deployment time. It was hell.
Over time I have come to prefer having two near copies that are each more concretely expressive of their task than a more abstract version that caters to both.
If you choose to not copy paste the code you better be damn sure the two places that use it are relying on the same concept, not just superficially similar code thats yet to diverge
Sometimes things are only the same temporarily and shouldn't be brought together.