Software Design Principles I Learned the Hard Way
read.engineerscodex.com
read.engineerscodex.com
Good domain modeling can enable making illegal states unrepresentable.
Sometimes, there needs to be more flexibility because of a lack of understanding.
However, other times you can put in a little effort and get the understanding needed for a good data model that simplifies future velocity and correctness a ton.
The classic OOP educational example of `class Dog extends Pet` is a case in point. In what sense is a dog like a pet? The only way to know the answer to that is to look at it through the lens of our requirements. Are we building a vet appointment booking app, or a video game?
Educational examples can be misleading due to lacking real requirements, and lead programmers to reproduce oversimplified structures.
[begin quote]
> [OOP] also make you have to think heavly about design, as you'll get yourself thinking about stupid things like "Does a buyer sell the goods, or does the customer buy the goods? What if we’re firing someone? Does a manager fire the person? Does the person fire themself? What if HR initiated the process, not the manager?"
This smells like very poor application of OOP. Let's take the firing example, it sounds like you're trying to design code like this:
employee1.fire(employee2);
But that's almost never what you want to do. Instead of designing classes that map to "things" like employees, it's likely your classes should be modelling business use cases like EmployeeTerminationProcess. To make up a bunch of stuff, let process = EmployeeTerminationProcess.create({
initiator: employee1,
toBeFired: employee2,
}).begin();
// later,
if (process.hasEnoughDocumentationBeenProvided()) {
mailer.send(process.makeTerminationEmail());
}
// or maybe,
if (process.isComplete()) {
process.archiveAllMaterials();
}
`create` might even be a factory which returns instances of different specialised classes for different processes. Say if the documentation requirements are different when the process is initiated by a manager or by HR.Trying to fit the colloquial English phrase "Bob was fired by Janet" directly into a line of code results in weirdness; instead of modelling English, you should be modelling the business process itself.
The problems you have with classes are common, and a result of crappy education and crappy existing codebases. It seems like industry and academia have gone through a wave of "let's model tangible objects as classes" which was never a good idea.
With any luck, the rise of functional programming will remind practitioners in OOP languages that abstractions are real (as David Deutsch described it) and that every tutorial containing class Dog extends Animal needs to be put in a bin.
[end quote]
[1]: https://www.reddit.com/r/node/comments/pyokmi/a_reflection_a...
On the other hand, fixing the messes left over by people with a penchant for overengineering keeps me rich. So have at it.
"Maintain one source of truth" is exactly what I mean by "DRY".
The question I ask myself constantly is "if this had to change, in how many places would I have to change it?". If the number is not 1, I consider it a "bug in waiting".
This is fairly simple for data. For code it can be more complex, but it's really about that same question. Is some algorithm changes, will I have to change code in 2 places?
I've seen people mistakenly bundle together similar code in the name of DRY like OP says, and it's an unfortunate thing. I recommend asking "the question" if you're unsure!
I rather copy paste my sql query and maintain 2 almost identical copies unless I am convinced they are actually a single query with some parameter.
Code which is necessarily the same might be something like “how to calculate an invoice total”. If you need to calculate totals in multiple places, you can’t calculate different totals for the same invoice. Have a single source of truth for this.
Code which is coincidentally the same might be something like “well, to get an invoice total we add up all these prices multiplied by quantity… and this spreadsheet we’re importing has the product cost and a multiplier to reach a desired profit, so that’s the same dollar amount multiplied by a number again!”. Creating a single implementation for this will only cause pain.
Even if these things are the same _right now_, there’s nothing saying they must remain the same. And as they diverge, you are going to be either creating brittle franken-code that causes bugs in seemingly-unrelated parts of the application when you make (slow, tedious) changes, or you’re going to keep only the repeated logic as it evolves and you will make the code more and more generic until it provides no useful abstraction whatsoever.
Don’t repeat things that must be the same. Do repeat things that just happen to be the same _right now_.
I'm a hobbyist. One of my apps recently had the exact same five lines of code in two functions.
After some thinking I decided to leave it like that. Moving it into a third function seemed just additional work without no benefit.
My question was: if I had to change A, would I have to change B? The answer was: Perhaps, but not in all cases. And if I had to, the change would have me look at both functions anyway, so it is not that I could just forget to change B. And it would be unlikely it changes at all in the next 2 years.
Agree with mocks also, currently not even using them, just go with integration tests and set db back to known state between every test run, oh and wiremock for rest
There's little to no point in doing mocks if your logic is not really complicated where you would need to separate testing it from the integration test.
I think a lot of people took "TDD" a bit too much to the heart and it ended up maligned where always more tests == better.
I know the book said so, but can we please just talk about the reality of our situation.
https://en.wikipedia.org/wiki/Rule_of_three_(computer_progra...
Rule of three ("Three strikes and you refactor") is a code refactoring rule of thumb to decide when similar pieces of code should be refactored to avoid duplication. It states that two instances of similar code do not require refactoring, but when similar code is used three times, it should be extracted into a new procedure.
If your “DRY” change PR removes 1 repeated piece of code but adds in 1 kind of nonsensical abstraction, 1 extra coupling between two pretty unrelated things that use that abstraction, 1 change to the input and output of the function, 1 test suite testing for two separate “things”, etc. it’s not sounding like such a clear positive contribution to the codebase.
If we want to go by some of the old adages then a non-dogmatic reading of Single Responsibility - the abstraction should be about one thing - is pretty good in my book.
Or I'm just dumb.
Now it feels like mostly gluing together cloud apis and fighting over which cloud product to choose and Managers who don't give a F about software engineering.
I’ve had a hard time explaining this but I know it when I see it. Two things that are the same shape but aren’t really semantically related. So if you try to share code, your abstraction breaks the moment they start to inevitably express how they’re not related.
It's never that cut and dry, there's always nuance to almost every pattern we try to apply in code.
The problem you're mentioning can happen regardless of whether someone repeats the same thing everywhere, or whether someone factors out common code into a single function.
John Carmack wrote an article [1] about the point you bring up, and how humans can understand a certain degree of complexity more than they can understand a very tiny amount of complexity or too much complexity. There is a sweet spot when it comes to function call length, not too long, not too short.
But once again this is a different matter altogether.
[1] http://number-none.com/blow/blog/programming/2014/09/26/carm...
Sometimes I avoided extracting a function out of some duplicated code because I intuitively saw that they had some potential divergence down the line, and it has paid off quite a few times. That's part of the problem as well, it depends on wisdom and intuition only gained through experience.
There could be some judgement calls and technical expertise needed to determine the right amount of parameterization of a function, but I'm not sure I follow the idea that a simple function, which is nothing more than a "name(parameters...)" that needs to be changed 6 months down the line is somehow causing a giant mess.
If anything, the benefit of deduplicating is that if you find there is a bug in your implementation, or an optimization, you can fix the bug or apply the optimization in one place and get the benefit at multiple call sites, as opposed to having to go to a bunch of places.
With that said, it's not my position to tell you not to repeat yourself. If you find repeating code helps you for whatever reason, go for it. I'm just saying that it's perfectly fine for developers to avoid repeating themselves by using functions and their code won't become a giant mess as you seem to indicate it will. Functions are very simple, flexible, and elegant abstractions for code reuse and even if you somehow found that you had to eliminate a function and repeat it in multiple places, that's a trivial operation to perform that most IDEs can do, so it's not like you're locked into this.
The same can not be said if you use a class as a way to avoid repetition.
For number 2, (please repeat yourself), the truth is in the middle. Don't repeat yourself when the proper abstractions present themselves, and always be on the lookout for a good abstraction. But don't force an abstraction where it's not appropriate. Where to draw the line is the hard part and in my opinion comes from experience and intuition. If in doubt, repeat yourself.
Number 3 (overused mocks) I find myself swimming against the tide of popular opinion. I think mocks are very valuable for unit testing. This doesn't mean you don't write your integration tests without the mocks, but I still think there's a ton of value in being able to quickly write a large number of fast unit tests that say "when this component does this, I expect the system under test to do that". As for leaking implementation details, I don't think it's such a big deal at the unit testing level. The details need to be tested somewhere. I appreciate this is an increasingly unpopular viewpoint but it has served me well.
For number 4 (minimize mutable state) this is software engineering 101, but always bears repeating.
Well, now you're tightly coupled against your single source of truth. That's a choice if you'd like your choices to be (a) more outages or (b) trying to build a perfectly available data store.
It's ironic that the example is banking; banks have many sources of truth and eventual consistency as patterns and always have.
Not only that but if I ever want to change data I now need to somehow inform the origin that the data has changed. Coupling me even more strongly.
The coupling is inherent. Things are simpler if you keep the source of truth in one place.
Regarding mocks, there are mocks that can hurt you and mocks that will help you. Mocks that can hurt you are mocks you generate by hand and it represents an idea of what production is, or maybe the untested, specified contract. Mocks that help you are the ones that are automatically generated representing actually what production has. Mocks like these have saved me from potential P1s many times - and millions in business losses. In an ideal world, I wouldn't need mocks, but also in an ideal world LLMs would do my testing.
I joined HN back in 2008 and I am mostly a lurker, but when I see an article that just plainly promotes bad practices with no samples of code or without going in depth. An article that would have been voted down or ignored to oblivion back in the early days of HN. I have to say something. I have seen enough bad code (specially in recent years) - I am afraid the new generation is getting their tips from all the wrong places.
I feel exactly the opposite. In an ideal world, I would write tests to define behaviors and LLMs would write and optimize the production code.
I'm a novice when it comes to mocking, can you explain more on this or link some articles to read?
Have you never seen more junior engineers make code 1000x worse (bordering on nonsensical), with them proudly saying they made it DRY-er? They are sometimes so focused on the one often-misapplied principle that they cannot reason about the produced code as a whole.
Also, the argument for 2 isn't very clear, it could be interpreted like the author doesn't understand inheritance.
Actually all four ways are possible: with either, neither, or both.
The GoF book says to prefer composition over inheritance much of the time, in an early page of the book.
Hear me out, I don't think classes should be used to mutate state, or perform actions on their own. But they are a nice way to namespace and think about components properly. It can be achieved with just structs and functions (strongly typed, for the love of all deities) but giving a container like a class has its advantages to namespace some categories of functions belonging to their own structs.
I agree that many other uses of classes (and fucking design patterns piled on top) were, and still are, a major source of pain in software development, I'm very glad to see that a lot of folks around the 15-25 years of experience realised the footguns are not worth it and adapted to write more readable "dumb" code.
Precisely. I’m a DBRE, not SWE, but I use Python a lot, and have written some internal tooling. Classes as namespacing is wonderful.
- A type
- A name for it
- A canonical way to build a value of that type
- A consistent way to reference its semantic meaning at runtime and in any static/build time/documentation format
- A singular, easily discoverable place to change it
You're gonna have to justify that. Can't just take your word for it. Where do you store your state, if not in objects?
OK, so in global variables, like in C language. And you think this is better than OOP. Glad that's working for you!
Even if there is some data I want to keep in memory there is no need for a global usually. In go I may just keep it in a var in a goroutine that doesn’t exit until the process does.
type Engine struct {
HorsePower int
}
func (e Engine) Start() {
fmt.Println("Engine is starting with", e.HorsePower, "horsepower.")
}Not at all!
This is the main reason OO has dominated for so long. OO has somehow taken on the meaning "Anything which isn't C".
And it's not even an accurate representation of C, it's a straw man version of C. C programmers know the perils of global variables.
Here is some code:
int myState;
void myFunc() {
}
If I take the above code and make it look like C: #ifndef MYSTUFF_H
#define MYSTUFF_H
int myState;
void myFunc() {
...
}
#endif
Yuck! Global variable! disgusting!Now I make it look like Java:
class MyStuff {
int myState;
void myFunc() {
...
}
}
This is somehow "encapsulated". But start calling myFunc() at different times from different threads and see if myState behaves more like a stack variable or a global variable.OO made it much more socially acceptable to pollute your code with this kind of "encapsulation". And OO devs decided that even this was too much of a straightjacket - some people liked C#'s take on properties, where you didn't have to work so hard writing getters and setters (further breaking encapsulation).
Real C Programmers™ actually encapsulate:
my_state_t myFunc(my_state_t in) {..}Once you take away inheritance ("prefer composition"), polymorphism (better offerings elsewhere), encapsulation (not encapsulation!), message-passing (better offerings elsewhere), you're left with a small syntactic trick that really wasn't that hard to do in C.
Even Lua gives you that [1]:
> This use of a self parameter is a central point in any object-oriented language. Most OO languages have this mechanism partly hidden from the programmer, so that she does not have to declare this parameter (although she still can use the name self or this inside a method). Lua can also hide this parameter, using the colon operator.
function Account.withdraw (self, v)
self.balance = self.balance - v
end
function Account:withdraw (v)
self.balance = self.balance - v
end
[1] https://www.lua.org/pil/16.htmlAccording to FP, data shouldnt be stored where possible and yes, if you can, store global or top level immutable blobs and pass by reference.
> That's how the real world is.
Well, maybe not. I dont rememer where it came to me that OOP is like classical physics, where everything has to get passed around with the speed of light and FP is like quantum physics, stateless until you do the computation. I dont know how accurate this analogy is but it still fascinates me since modelling reality is enlarge what programming is.
Structs and functions are great if you see yourself adding more new functions than data types in the future (the existing functions don't need to change!). OOP is great if you see yourself adding more data types (the existing classes don't need to change!).
Better to know how much time things take.
I'm mostly a web dev, but I was experimenting with some very mathy Rust code. I did some back of the envelope calculations about how many cycles a CPU might take on various floating point operations, multiplied that by the number of items I was pumping through it, and arrived at a number that was 4x out... then realised the Rust compiler had performed auto-vectorisation with SSE2!
I wasn't spot-on, but I was within like 20-30% I think. And that made me feel like a real engineer haha.
> but it’s important to figure out what is truly necessary storage-wise versus what can be derived on-the-fly. In the “v1” of something, I’ve found that minimizing as much mutable state as possible gets you pretty far.
It seems like the author is specifically calling out that you shouldn't store derived data unless you either can't recreate it easily or have measured (usually a post-v1 activity) that it is faster to store/cache that derived data.
I'm triggered by context free "I usually find that..." and "In my experience..." assertions.
DRY != god class / god methods. Strawman when a common problem is trivially-duplicated code sprawls and spews garbage and poorer maintainability over a codebase.