Code doesn’t have to be a mess
danielsieger.com
danielsieger.com
People abstract before an abstraction is necessary.
I find single file dense leetcode style code easier to understand and follow the flow. Algorithmic code I can reason around. A large mature codebase is far harder to get to know.
One of the first things I do when I study a new codebase is find all the entry points and follow the flow of code from beginning to the thing I am interested in.
One person's beauty is another person's mess.
It's harder to change an existing codebase than to write a simple program that does the new thing but not in the context of the original program. A reference implementation of the various components is far easier to understand than one big ball of mud. Fitting problems together is hard. You need to understand the old thing before you can introduce the new thing and it ends up being forced or hacked in if the design doesn't support the new thing.
I tend to write reference implementations of everything, then combine them together as a separate project.
I find an empty file far more reassuring than a large codebase.
Refactoring will give them the chance to see what the actually moving parts of code are.
All the files and code on the way to get this to happen such as Classes, parameters, arguments, variables, functions, methods, closures, objects are ideas of the languages compiler to abstract the instruction stream.
Command line arguments, class constructors, URL query parameters, marshalling, JSON field names, method parameters, function arguments, HTTP headers, cookies, request objects, events are just complicated variations of passing data in the right shape. They are not the above list of "moving parts" or computation that is easy. In other words modern coding is just configuration.
I feel the complexity of modern code is a problem we created. And I feel there's something missing. It's hard to update code.
When I find the loop that does the thing, I feel I can understand the codebase such as the magic +1, -1 or the relationship of objects linked together in a data structure or the assignment to a list or array or variable.
"How does that get to here"
Where I'm definitely weird is that I have a higher verbal score than your typical developer, and I'm not afraid to use a thesaurus to find a better word for something. Too often we end up recycling jargon in situations where they are not quite doing the same thing but nobody could be arsed to open thesaurus.com and find a word that telegraphs, "B is like A but is not actually A."
This is why Sum Types are so great. They give you an option between "concrete type" and "any type that implements this interface": "one of these specially enumerated types".
I like the way you used it. I'll try to use that myself.
Them (on the subject of 4 new methods): hey why did you make method 3 look "weird"? You should make it look like the other three.
Me: because one of these methods can set the building on fire, and the rest can't. I bet you can guess which one is the dangerous one. Works as expected.
Overly generic code is often pre-emptive, and most times the day never comes that you need that flexibility. And often when you do, you discover you need flexibility along a different axis anyway.
And most of what we do is story telling: what were the requirements we understood? What is our model to solve it? What are some precise examples that show it working in different capacities? When people treat the code as simply "the thing the computer interprets", instead of "the thing the next person has to comprehend", you get this inevitable slide into incomprehensible code.
Unfortunately our profession is obsessed with outdated approaches to (premature) performance optimization, and an addiction to being "clever".
I think I'd much rather surround myself with driven folks that have empathy and a strong desire to be understood. That's a long way from the programmer stereotype I'm familiar with.
More tightly bound code is often easier to understand and mechanically modify later - there are fewer places where you lose "if it compiles, it works" guarantees.
I feel like a lot of people are blindly pulling coding habits from libraries, and applying them everywhere. Libraries and applications (i.e. "terminal" products not used as a library by someone else) have different needs and different goals - don't write your application like a library, it'll be a huge pain.
This!
I genuinely can't tell if you're being serious or not. If you are, do you also like to read books written as one giant chapter? Or entire chapters as one giant paragraph?
Without documentation I find large projects difficult to understand. There's literally too many global symbols and I cannot see the forest for the trees.
What's the model of this program? What are the core principles that the author is using? Do I really need to read every file to understand what is going on?
With Leetcode style programs there's one file with everything in it and I can usually find the entry point. The problem is well defined.
I can see the moving parts in a Leetcode style problem. The looping, the data structure creation and control flow, arrays and recursion.
Large mature codebases such as Java projects have thousands to millions of small files and packages it can be difficult to see how things fit together, every file seems to be 10-20 lines long.
I like C projects as they have lots of code in one file. Everything I need to understand a module is in one file and I can use vim folding.
Splitting code into multiple files really only makes sense to me when there’s a clear division of responsibilities. That can mean a lot of things - like client / server, utility methods / core algorithm or class A / class B. But plenty of complex data structures are much easier to read and understand all at once. For example, I have a rust rope library which implements a skip list of gap buffers. The skip list is one (big) file. The gap buffer is another file. Easy.
Good abstractions let you see the shape if the forest.
Most "good code" or "simple code" is an inedible potluck of "simplest solution for the problem at the time" with some documentation.
Actual good code has intention revealing abstractions that communicate the essence of the problem and problem domain.
Not sure about literature but for almost any technical matter I prefer learning from the smaller details instead of the big picture - it's often too vague and just doesn't stick in my memory.
1. read one line of first chapter,
2. then skip to the last sentence of the middle chapter,
3. then realize the first chapter was actually the penultimate,
4. then read the forth sentence of the first paragraph of the second chapter, bearing in mind what you have learned,
5. then throw your hands up in the air in dispair
For a maintenance programmer, they may already understand how everything works. They aren't following a particular code path, necessarily. Maybe they're working on a new feature and they need to re-familiarize themselves with previous chapters. It that case, it's nice to jump to a file that concerns itself with things grouped together.
fixed it for ya :)
I'm a maintenance programmer. Even working layers of layers of layers above the actual shit does not make it not stink.
>> jumping around
...should be intuitive and joyful, not a disaster to your brain.
EDIT: I am a fan, though, of SOC. I guess enterprisey code tries to be that (but fails hard at it).
Is this you showboating the greats of layered code?
Depth is not width and width is not depth, but surely you see the difference in the two?
Which is not at all like a modern codebase that is modular, abstracted, etc. If you're new to a codebase, and want to understand one particular feature, you'd likely need to jump back and forth across 10 files.
It's not far-fetched to say that makes it difficult to understand.
Let's say you have a simple endpoint that takes a list of comma separated inputs, parses them as numbers and spits back a sorted version of that.
I don't want to see a version of that, which a compiler might have inlined. Including the implementation of the sorting algorithm. I only want to see a high level of abstraction version of it. Basically just something like (pseudo code in a non existent language) :
fun endpoint(input):
inputs[] = split(input, ',')
numbers[] = parseAsIntegers(inputs)
return quicksort(numbers)
I can easily understand what this does and what the idea is behind this "algorithm" in 3 lines. If I had the "inlined" version of this I would have to manually identify each of these parts and potentially skip over tens to hundreds of lines.I think this is really a bit about trust. Do you trust that these named functions I am calling do what their name does? Does quicksort actually do a quicksort or has someone implemented bubble sort in there? Of course this is a minimal example and especially quicksort would probavly just be a library but imagine all of these were large complex pieces of our code base.
Personally I am an advocate for using functions (methods or whatever your language calls them etc) and naming them properly and then trusting those names by default. I want to spend time making this nice and understandable and abstracted once when writing. Not every time someone reads it. As soon as something does not seem to behave in the way the name suggests I will then and only then go check the actual implementation and for example find out that parseAsIntegers actually also supports floats and quicksort is not actually quicksort but bubblesort and that is why this endpoint was slow etc.
This discussion often ends up in the extremes: inline everything versus abstract everything. I don't think anybody reasonable would opt for either of those extremes, we should focus on the very large middle ground where there's a lot of subjectivity.
Trust me on this, I've been raised on the DRY dogma and all related architectural patterns in favor of abstraction. I've lived the life, for 2 decades. But I cannot ignore the outcomes. Most codebases are extremely difficult to understand and it's very painful to change things. As 90% of all software development is maintenance, that's a planetary-sized problem.
This doesn't mean you should inline everything, it means sane choices. As a simple example, say you're using a literal in your code:
(if orderAmount > 100000)
This code is incredibly easy to read. Common convention says to put this in a constant, at the top of the file. Old me agrees, new me does not. For as long as it's the only occurrence of the value, it doesn't need abstraction. The only thing it would do is make the code more difficult to understand. The very eager abstracter might even put that constant in a separate file.
The point of this example is to abstract based on real reuse, not imagined reuse. I'm not against abstraction only against unnecessary abstraction.
A second example. Say you have a reusable UI component. A change request comes in that's pretty large and specific for one niche need. It kind of goes against the spirit of the initial purpose of the component but is still related enough to consider it in scope of the component.
Old me might add a "toggle" to the component, after which it can render in two modes. Sometimes called a "god component". This approach sucks. It makes the component much more complicated and changing and testing it becomes a nightmare.
Even older me would break down the component into smaller components and then "compose" them based on its mode. This is even worse, now you have to jump around many places whilst the sub components are never actually reused (pointless abstraction).
New me says fuck it and splits the component in two. Allowing significant code duplication between both components. It's not as radical as it sounds, it's in fact incredibly comforting. Each component can easily be understood (less complexity) and making changes becomes far less stressful as your blast radius is tiny.
Developers spent the vast majority of their time not coding, instead figuring out how something works and how to make a change that doesn't break anything.
What I do have to have to question is the strict non-use of constants. It can be very very useful to use constants for such things, e.g. if you are calling libraries that do not make it apparent what is what. Say you have something that takes a timeout value.
send(data, 10)
What is this? I have to know what send is, what parameters it takes etc. I might have to look that up. I can easily work around that with a constant. const timeoutInMillis = 10
send(data, timeoutInMillis)
The same principle can easily apply for other similar situations. I really like it for things like doSomethingThatCouldTakeLongButAlsoShouldHaveATimeout(data, 64800000)
What is that and what does that value even mean in human readable? Of course some of these values you will recognize if used enough but so far my domains have been sparse enough that I don't recognize all of them and have to compute. Much better (with shorter, real names anyway but ya know, we're dealing in simple examples here :) ): REALLY_LONG_RUNNING_PROCESS_TIMEOUT_IN_MILLIS = 1000 * 60 * 60 * 18
doSomethingThatCouldTakeLongButAlsoShouldHaveATimeout(data, REALLY_LONG_RUNNING_PROCESS_TIMEOUT_IN_MILLIS)
So far most people I've talked to find that it's much easier to recognize that this has a timeout of 18 hours but the method happens to want milliseconds.Oh and don't get me started on people that use the constants from the real code in their tests, completely defeating the testing. Especially if they then do math with the constants and simply copy the math - or worse, put the math into a method and call it from the tests too - to their tests. Test expectations have to be computed once, when writing the test and just hardcoded into them, otherwise they serve no purpose as changing the code itself will always result in green tests even if you've just made a major mistake by changing the values without thinking.
public class EndpointManager {
private EndpointInputManager eim;
private StringSplitter splitter;
private NumberParser parser;
private Sorter sorter;
public EndpointManager(EndpointInput input) {
eim = new EndpointInputManagerFactory().setInput(input).build();
splitter = new StringSplitterFactory().setDelimiter(new Delimiter(",")).build();
parser = new NumberParserFactory().setFormat(NumberParserFormat.INTEGER).setMode(NumberParserMode.LIST).build();
sorter = new SorterFactory().setSortOrder(SortOrder.ASCENDING).setAlgorithm(SortingAlgorithm.QUICK_SORT).build();
}
public EndpointOutput endpoint() throws ParseException {
splitter.split(eim.getInput());
parser.parse(splitter.getList());
sorter.sort(parser.getOutput());
return new EndpointOutputFactory().setOutput(sorter.getSortedList()).build();
}
} @Path("/sort")
public class SortResource {
@GET
public List sort(@Body String input) {
validate(input);
return Arrays.stream(input.split(","))
.map(Integer::valueOf)
.sort()
.collect(Collectors.toList());
}
}
I do recognize the kind of code you pasted. Had to work in code bases like that for way too long. Never want to work in one of those again. There's probably lots of EJBs and other such nonsense around that?This one really frustrates me. Write code to the complexity level needed to solve the problem, and nothing more. The only time I'd break from this is if I know for certain that the added complexity is going to be necessary in the near term.
> ... not all refactorings improve the code.
While true, I have a low tolerance for code that requires constant bug fixing, or is so overly complex that the thought of modifying it makes you want to cry. Some projects truly require that level of complexity. But in my experience many do not, and once you've gained a solid understanding of the problems it is trying to solve, incremental refactoring is a fantastic way to improve the code's stability and maintainability. This is especially true in C++.
Mastery will be, when you write code in a way, that does not impose unwarranted limitations from the start, and still keep it readable and only containing mandatory complexity.
Usually this can be achieved through deep understanding of the problem, mapping to simple concepts or finding or making that one concept that captures things well.
Not always it can be done. Not always can a masterful solution be found, which keeps complexity low. However, it is definitely a mistake to draw a black and white picture of "if you want to make it work for the future, you must add complexity". Often people simply choose bad abstractions or wrong ones and will only realize, when the future has become the present and the system they built cannot fulfill some requirement.
This is me. They tend to lead me to good abstractions (after merciless refactoring), and I'd like to think that I'm sucking less at this over time. But my overall process is very slow (good thing I'm self employed). Understood that it'd be better to stop and think instead of diving into new-abstraction boilerplate work.
Worse though is to be under heavy pressure to ship and move on — with the bad abstractions getting hopelessly calcified / buried.
When writing software, ideally I'd like to make all the right choices and use simple implementations of abstractions that do not impose unwarranted limitations.
I think that it's sometimes worth it, early on, to do things the quick way despite bad abstractions. This can get you to a place where it's easier to reason about good abstractions.
Sadly, I've been on teams where a bad abstraction was adopted because it was just assumed that that would be quicker. Instead of doing it the quick way, we just did it the bad way.
> This one really frustrates me. Write code to the complexity level needed to solve the problem, and nothing more. The only time I'd break from this is if I know for certain that the added complexity is going to be necessary in the near term.
No, the ecosystem the code exists in and my ability to reason about the codebase is worth way more than any gain that comes from blindly stacking "simplest solution for problem a, b, .. z" atop one another without regard for higher level understanding of a codebase.
> This one really frustrates me. Write code to the complexity level needed to solve the problem, and nothing more. The only time I'd break from this is if I know for certain that the added complexity is going to be necessary in the near term.
I worked with a guy who did that. He had a plan for what the project would look like 5 years down the road, and he built abstractions to support that. He could get away with it because he could hold it all in his head and it all made sense to him. When version 1 was half finished he was called away to work on another project, and those of us who followed in his wake struggled to make any sense of what he left behind. A year later he was laid off. The project was a success, but nobody ever asked for version 2.
This is so true. Also I believe that when the original author wrote the code he had a (hopefully) clear vision of the solution. He wrote it as tidy and fitting to the problem as he saw it. Then sometime later someone else comes in and is supposed to alter the code in a way which does not fit the original author’s idea of the problem. This creates a mismatch. The new guy can’t and won’t change the code too much, as it is too risky/much to do and therefore will only do as little as possible to make his change. Then some other comes along and do some more changes, which again isn’t enough etc etc, et voilà, you have a ball of mud that screams of a rewrite.
This is how I write software. My stackblitz is full of domain independent experiments https://stackblitz.com/@Pyrolistical. I then copypasta this into private projects once I figured out how the individual piece works.
That's not a criticism at all by the way, I just found it striking.
In hindsight if you focus on three questions: is it good? Is it right? Is it true? You'll head toward a good direction
Synthesis of ideas is really important and the building blocks of understanding are fascinating. Programming and mathematics is taught as building blocks and then deliberate practice.
I really enjoy reading plain descriptions of things, especially of other people's code.
If you understand the core insight, difficult things can be easier to understand and apply for you.
I really want to understand how tracing compilers work and LuaJIT, JVM and V8 but I found the code a bit too hard to understand as I jumped into the wrong locations.
There has been two instances where Wikipedia was enough for me to understand and write an algorithm that implemented the description. Wikipedia doesn't have pseudocode for multiversion concurrency control but it does have an accurate if subtle description. I did the same for btrees but I did read some other people's implementations to get a feel. I of course wrote mine completely differently.
I want people to document their code enough so that the core principles or idea behind their code could be reimplemented by someone else just by reading the description of how it works.
Rpython and Pypy documentation is good but I still don't understand it enough to implement what it does. Which means I'm missing some detail or core insight.
Same for rewrites. Often the rewrite will have the same number of problems, just different ones.
Refactoring was absolutely necessary for this. He was writing a single simulation program that was single file in size. But once he'd created ten subroutines all changing a raft of the global variables, the slightest changes produced hair-raising bugs that he'd obsessively dive into debugging.
The intuitions of structured programming and object oriented are more important than absolute fidelity. My points were: "If you can't have an object here, at least have a well defined, standard interface to values that need to be in a consistent state" and "decompose long action sequences into subroutines and if you can't do that, least group similar actions with similar actions in that long action sequence".
Which is to say a given piece of code might not the structure you want but if has a structure, that can be enough. But then again, that piece of code might not structure at all and then rewriting it really is necessary and often is easier than debugging it a few time.
And working with a large piece of "bad" corporate code, I've more than once that you something with one sensible if idiosyncratic structure that was refactored more than once by people who didn't understand the structure and imposed their own structure on just part of the code. But through an exercise in archeology, one can make the whole artifact work.
But that doesn't mean you can't have code that is a true mess when the writer has no experience and no concern with structure.
> I tend to write reference implementations of everything, then combine them together as a separate project.
I find it really difficult to go through huge chunks of iterative code. I need abstractions otherwise I can't get my head around it. I often wind up refactoring into manageable chunks (even in pseudocode/diagrams) just so I can understand stuff.
For reference, my cognitive abilities are heavily skewed towards verbal/abstract reasoning - like several standard deviations above the norm - and my spatial/concrete reasoning is nearly the inverse of this, it's terrible.
I wonder if this has something to do with it!
Understanding lots of code at a module, function, or even more granular level with a magnifying glass feels more productive than struggling to understand the full picture.
It also rewards you with instant gratification. Reading and writing ncrete code gives much more immediate gratification.
Sometimes an abstraction cuts to the core of the reason why.
See for example https://algebradriven.design/
Good abstractions can communicate intent better than mounds of concrete code because they speak at a higher level.
However, mounds of okay concrete code is way easier to deal with then poorly thought out abstractions.
This means pragmatists get little practice in abstractions, where their pragmatism is needed most to uncover the useful abstraction and avoid the overly complex invented abstraction.
Abstract code also has the advantage of parametricity in strongly typed programming languages.
https://GitHub.com/samsquire/algebralang
It's designed to be expressive and powerful and practical.
The core insight to a problem is rarely what we spend most of our programming time doing.
I believe that is a mistake.
I'll have to check your language out though!
That's a great point. I once read (here on HN I think) that the value behind a piece of software is not the code but the team whose members all have the same mental model of the problem and can successfully map it to the code. Lose the team and you lose that map.
> One of the first things I do when I study a new codebase is find all the entry points
I follow the same strategy but… Good luck with that when you're facing a Spring application. :)
The first step is getting the lexicon right. Frequently the business lexicon is ambiguous in such a pervasive way that the people immersed in the business aren't aware of the discrepancies. For example I remember from working in healthcare the words "claim" and "member" often have very different meanings in different contexts and I would see developers hacking code together to get the data model of one context conform to the data model of another when they should have been treated as different entities.
Getting requirements in writing was an uphill battle and the lack of requirements always wound up screwing over the developers because there was no contract to prevent scope creep and the developers were the ones that were held accountable for misunderstood features and missed deadlines. As a result everything was constantly rushed and not well thought out. It took me a long time to convince my boss that the issue stemmed from unwritten requirements and a lack of planning.
To add, how could they do any QA when testing needs to map to those unwritten requirements
Getting junior devs to do this is like pulling teeth. Trying to get a feature stopped after they've built it is soul crushing for them. It's a problem.
At this point I've all but given up beyond minimizing the blast radius in code review.
Perhaps tech companies could have kickoffs/workshops where the participants would create sand mandalas together?
The guy who just poured the concreted for a foundation, does he care that much that it's torn up or not use? Probably not, even though he's likely skilled and professional.
We are far too precious.
Beyond opportunity cost, you can think of it as deleterious to your performance. If 10% of your work never gets merged because of shifting priorities, compared with someone else who has miraculously dodged these problems, that's pretty unfair.
Clickable wireframes, design sessions, mockups, etc - they can all help explore ideas before code, and potentially save things. I've had numerous examples where I can identify "this is confusing" or "this doesn't solve the problem, just moves it around a bit" and I'm usually 'outvoted' by others, and do the work. It's usually only after it's in peoples' hands that they identify the rough edges (or more).
That was a frustrating period.
Hilarious. I wonder what a plumber or carpenter would say if you were to complain to them on how unfair your job is, because 10% of your output doesn't show in the finished product, yet you are still paid for that output. Imagining the reaction to that just made my day.
I feel like this is a pretty uncharitable caricature of my position. I'm not whining about all my work not getting in. I'm saying, "We're building houses for poor people, I care about this, you had me spend a week on this thing you said was going into one of the houses, you were wrong, I blew a week of work that could've gone to building the houses, and winter is coming". It's not about me personally, it's about me caring about efficiency.
> yet you are still paid for that output
I don't only work to be paid. It's a necessary but not sufficient component. I try to find fulfilling work that I think improves peoples' lives, and I'm fortunate enough to achieve that more often than not. I'm not saying I'm not selfish, just that I'm not entirely selfish, haha.
It's a worldview that's so nonsensical that it's its own reductio ad absurdum: If a fast-food worker complained about wanting to be treated with dignity at work, would you similarly scoff at them because coal miners or sweatshop workers don't even get physical safety?
Having high standards is a _good_ thing. It's the hallmark of society's progress. It doesn't preclude being grateful for the privileges you do have, and it's nothing to be ashamed of unless your self-esteem is so low that you think you don't deserve to be treated well.
> People's expectations are relative to their environment. [...] Similarly, the vast majority of HN users are extremely high-percentile for global wealth and income
> Having high standards is a _good_ thing. It's the hallmark of society's progress.
seem to suggest you subscribe to the "trickle-down" ideology. I don't. No, having pockets of "extremely high-percentile" people who are entitled to complain about "unfairness of 10% of their work not being appreciated" is not a hallmark of society progress. It's closer to systemic exploitation. It's a pattern we should know very well from history lessons. No bread? Let them eat cake! Sure. Just brace for the impact when the bubble bursts - there's a sharp blade at the end of this road.
> it's nothing to be ashamed of unless your self-esteem is so low that you think you don't deserve to be treated well.
Because having 100% of someone's work accepted as useful when it's not - for whatever reason - is a basic human right that everyone deserves. That's called "being treated well". I didn't know; I thought not getting 100% sunny days in a year is called "just life", but now I know it's a violation of my rights. How could I be so wrong for so long?
I'm being sarcastic, but you have to accept this comment in its entirety and tell me how happy you are that I wrote it for you. I put work into writing it. I deserve being praised for it, no matter how much you like what I wrote. Right? Please, do treat me well.
How's that for a reductio ad absurdum?
Again this wasn't what I was saying. My argument is that engineering time (and time in general) is valuable, and we should be careful how we spend it. In other words: if, over the course of a project, you're wasting a lot of time, that's bad if you care about the project (and you probably should). Maybe you disagree, but I don't think this falls under the "entitled millennial SWE" category, but rather the "we can do better" category.
I'll also, for the sake of discussion here, say I've done some pretty shitty work and have benefited tremendously from code review and general discussions with my colleagues. I'm definitely not someone who thinks they're a "extremely high-percentile" person, probably above average, but definitely not like a Brad Fitzgerald or something.
Myself I not too long ago did a 6 month crunch with the rest of my team on a product that was cancelled right before launch...
Of course, you and your team should work together to _avoid_ having to reject work! But it can and will happen, it's perfectly normal for mistakes to be made, it's how all humans learn.
Trying to deny that people sometimes fail is foolish. Punishing yourself for making a mistake is on you.
The only unfair thing here is taking others in the team hostage with the idea that you are entitled to getting your work merged regardless of its quality, purely because it would make you peeved, cranky, annoyed! That constitutes toxic behavior. If this is a pattern for you, people will avoid working with you.
Instead: embrace the opportunity to learn. Get feedback, reflect with the team, do better next time. Maybe pair up to refactor your work. Take the positive approach!
- Someone says "build this thing"
- I build "this thing"
- That someone says "just kidding, we're not gonna use it"
- I'm peeved
Someone else in this thread is arguing this is an entitled position, and here you're arguing that... well, I think you're arguing that I think all my code is always amazing and should always be merged.
I'm not! Like I wrote elsewhere I've written some pretty shit code, I've built the wrong thing, and I've built broken things. I'm sure this is true for most SWEs. This isn't the scenario I'm describing.
But I think this discussion has some merit in terms of how we navigate code review. For example, conversely, I've been on the other end of some pretty... bad feedback. The first example that comes to mind is that we had a portal where you could search by text or category, but once you selected a result we wouldn't save your search anywhere (query params, session storage, etc.). Consequently, when you clicked our "back" link, your search would be gone. We YAGNI'd it for a long time, but we accepted a very tight deadline project (COVID/government related) that required a category that needed to be sticky.
I built this using query params, like pretty much every search out there (for good reason). This ended up changing a lot of templates, a couple of front-end React components, and required extra logic in a couple Django controllers. It was a big-ish change, maybe (to my recollection) 300-400 lines across a few stacked PRs--meticulously, for ease of review. All previous tests passed, all the new tests I wrote (typically >= 50% of my PRs were new tests) passed, I even built a punch list of UI tests I ran through (this was going to be a big user-facing feature and I wanted it to be bulletproof). This took I think... 2 days of constant work, so something like ~30 hours.
This wasn't our typical process; we skipped our usual engineering meetings about implementation strategy and what-not. Our team was small--4 people including our CTO--but even so we had a wide diversity of opinion when it came to implementation, architecture, and style, so it kind of ended up being the case that if we wanted anything to get through PR we had to hash it out beforehand. But we literally had 7 days or something to do this, so we just didn't have time.
But, predictably, despite all my tests and punch list, my PRs were rejected as "too much code", and we missed our deadline. We launched without the feature. Our CTO reviewed the vast majority of our PRs, he reviewed these and he was pretty furious about the scope of the changes, blaming me for missing the deadline.
Afterwards, he tried reimplementing it using query params in less code, but failed. He then tried reimplementing it using local storage, which was less code, but had multiple problems: local storage works across tabs which is deeply weird, but even if he switched to session storage, it didn't work in lots of versions of mobile Safari if you're in private browsing mode. I rejected that PR for those reasons, which we disagreed vehemently about. Eventually, a couple months later, we paired on it, and basically reimplemented my work together.
There are obviously a lot of flags in this little story, but I don't think they're wildly out of the ordinary for a startup (if anything, it's way too much process for a 10 person company). My point is that, while I'm sure there are a lot of cases of "I'm God's gift to this company merge all my work never question me" out there, there are also a lot of cases of "no PR is fit to merge the first time" and "I'm a great programmer, you didn't do this the way I would, therefore this isn't good enough" as well.
Is there no planning flow where you work/have worked? I understand giving developers freedom to build, but some sort of oversight by someone with a view of everything that's going on is also necessary.
Typically there are entire planning exercises that happen before stuff even hits a JIRA board.
These are usually conducted by Product Managers, Project Managers, analysts, and often in startups the CEO themselves.
The fact that people are off building features willy nilly sounds like it would contribute to messy code base.
Knowing what's next helps plan the work on the developer side as well which also minimizes the blast radius.
That's why I suggest first filing an issue, discuss a design, and only then actually implement the feature. Will save you many hours.
As an example, I did an spike to explore a request from another team. I wrote a short document with my findings and recommended against it due to the cost/benefit analysis. My manager told me not to use this story in my performance review since "we don't want to display failures".
We’re amidst a rewrite and we have an off-shore team involved. They same team that built the original starting 4yrs ago.
One of their team members decides “we should validate the TLD for email addresses entered by the user.” Code is added, a TLD file is added, and in code review I reject the whole concept. Show me the ticket or feature docs, and I’ll argue with the author of those instead.
“We did this in the last app…” Maybe, but we didn’t spec that for this app.
He got his local project lead (non-tech) to write a Jira story for us to “discuss the technical implementation.” Dude, srsly.
Our app is web-first and is used on congested mobile networks (like, hundreds of people all using the same cellular site simultaneously.) A TLD file does not need to be delivered to each of them for validation that’s pointless.
The idea is off the table, code rejected, but the guy spent time doing something no one asked for and had his local team onboard with it.
It was an eye-opening experience.
My style is influenced by Haskell and Rust, even when I program in, say, C# or PowerShell.
A simple example: I will extract the read-only logic into a pure function and minimise the size of the mutable procedure. This makes it trivial to test the logic in isolation without triggering any side effects. Similarly, the logic can have convoluted control flow but the imperative code can then wrap that with a single try-catch block, transaction, or retry loop.
For me this was such second nature that I didn’t even realise I was doing it until I saw the imperative spaghetti written by the juniors. I tried to explain with pair programming sessions what the benefits are of my approach.
Without fail, they would just “hack something” into the existing spaghetti, adding yet another mutable global variable to track some new state.
In every case they said they were in a hurry and that they would “fix it later”.
I replied: “there is no later.”
This. We need to stop using time pressure as an excuse to do a bad job. Moving things out to a separate function might add 10 min but save 100x when everything comes crashing down.
Of course you shouldn’t overengineer but so much can be gained from spending a little time just thinking about how this should work.
It's like a conservation rule in physics, for every bug found, a bug fix must eventually be implemented. Bug in, fix out.
But... if you leave a bug lingering, then it can cause test failures for unrelated code development. It can trip up other developers. It can cause false positives until resolved.
So the only logical conclusion is that all bugs must be fixed ASAP, otherwise they have a "multiplier" factor dependent on how long they're allowed to persist. If left unchecked, this can blow out exponentially, until you're unable to efficiently fix bugs because you're tripping over thousands of other unfixed bugs while doing so.
You would think this kind of thing is logical, but no-one ever believes me. There's just slow blinking and then a slower repeat of the same old mantra: "We'll fix it... later?"
Over decades I have compiled my own list which contains all these and bunch of other behaviours that are needed for successful project.
I would add one or two very important thing missing from the list.
One, not explicitly mentioned but covered in other points is to plan for simplicity. Make simplicity an explicit goal of the project and set up process to remind of it at various important points in the process. For example, I have a checklist for adding a new technology which has a very long list of things you have to think about before adding new tech of any kind (like "is it possible to replicate it with couple pages of code"). My goal is to have tech stack so simple that newcomers can feel right at home and productive immediately.
Even if I (we, me and my team, whatever) screw up, then the future owner will tend to have much easier time fixing it if we tried to keep it simple. My hardest challenges were not difficult technical problems (most backend applications tend to be very simple problem from technical point of view) but rather past teams that were very smart and created a monster so complex they themselves ground to a halt after some key people left.
And connected to it (part of the checklist) is to be aware of when you are about add things to the project for intellectual gratification rather than practical purpose -- and cut it mercilessly out.
Engineers tend to really dislike working same technology all over again, but this is what is needed to become really proficient. It seems exciting, but every time you add or change something in the stack you need to learn that thing (and accept being less productive for some time), you accept risk of new problems (and risk is a cost) and, finally, you cause the same to every team member and any future hire.
And while it is easy to see the benefits of something, the costs and risks are usually much less understood before you have invested enough in it. And, additionally, frequently the benefits are much overvalued.
I especially like your point about planning for simplicity and making it an explicit design coal. Communicating this clearly to the team seems super important.
As for adding things out of intellectual gratification: This is so true. I've seen this all too often, myself included. Many good engineers are curious by nature, and it can be tough to restrict this curiosity. Maybe it's a question of having different outlets for that sort of creativity, either at work or in private.
If your list is available in some form, I'd be curious to have a look!
The job of the most experienced person in the software development organisation should be to spot unnecessary complexities and find ways to eliminate them.
The issue is, most experienced people tend to be engaged in activities for political reasons like adding new technology -- which tends to be perceived as more valuable than removing it.
I might be biased (by selection). I am called upon to join and help projects that face significant problems (emergencies all the time, no time to breathe, way behind schedules, unable to deliver anything, etc.) But every time I join a project like that, the repairing process tends to start with removing stuff rather than adding. And if stuff needs to be added this is usually so that it makes possible to remove much, much more of complexity somewhere.
As an example, we have stabilised at least 3 projects by removing "microservices" and rolling everything to a single monolithic application. Not saying microservices is a wrong idea, but saying it might be wrong for a particular project without strong automation culture, tooling and without large enough problem to solve. Somehow this always starts with strong opposition and ends with happy people that can code and not spend significant portion of their time dealing with complex, flaky infrastructure.
My rule as I present it to the team is "I want at most one of anything unless we understand exactly why we need more than one." So one programming language (unless you need one for backend and one for frontend, then we need two), one application (unless you have multiple teams and then one per team might be better to make them independent), one repository, one cloud infrastructure provider (AWS tries to be on parity with GCP, why do you need something from GCP just for it being incrementally better?), one place to store all documentation and procedures, one database tech (do you really need half of the application use MongoDB and another half use Postgres?), etc.
The rule might sound childish, but it is simple and helps people make better decisions on their own which is essentially what you as a tech lead want.
Which altitude of simplicity is most valuable?
`man ssh` gives me detailed descriptions of all flags.
'UNIX Style, or cat -v Considered Harmful' http://harmful.cat-v.org/cat-v/
Is this a typo or a proposed Bourbakism?
BASH_BUILTINS(1) General Commands Manual BASH_BUILTINS(1)
NAME bash, :, ., [, alias, bg, bind, break, builtin, caller, cd, command, compgen, complete, compopt, continue, declare, dirs, disown, echo, enable, eval, exec, exit, export, false, fc, fg, getopts, hash, help, history, jobs, kill, let, local, logout, mapfile, popd, printf, pushd, pwd, read, readonly, return, set, shift, shopt, source, suspend, test, times, trap, true, type, typeset, ulimit, umask, unalias, unset, wait - bash built-in commands, see bash(1)
First, people generally don't hire me to set up a WordPress site (I should get into this though); they hire me to write something new and bespoke. So my skills are in exactly that: I build new stuff.
Second, I'm pretty bored by the idea of gluing dependencies together. It's neat to see how fast or neatly I can do it, but that's good for a month or two tops.
So if you want me to cook up new tech with a pretty good amount of code, I'm your guy. If you want me to carefully build something someone else has done 100x before while constantly having meetings about capitalization, line length, and coding-fad-of-the-week stuff, I can't handle it. My (totally rational, at least to me) response will be: customize a CMS for $10k, don't hire an engineering team for ~$500k.
If I'm stuck on this project, I'll subconsciously try to introduce joy into my life by doing bad stuff, like writing a lot of cool new code where I shouldn't, and so on. Our incentives are misaligned.
---
Or, you can think of it in terms of innovation tokens. Are you building a new database storage engine? Adopt the conventions of the database you're building it in; don't also try to innovate a new architecture/style. Are you building a new JS framework? All the innovation there is in developer experience, so all your innovation should go into abstractions and mental models; don't also include new, surprising algorithms.
By picking a thing you are doing, be aware that there are 10000 other things you picking to not do.
Obviously we don't want a complete dumpster fire of a codebase, but some mess is inevitable and healthy. First see the mess, then refactor. Refactoring before the mess is how you end up with crappy abstractions.
In the past few years I've adopted the attitude that code cleanliness isn't really that big of a deal. There are some obvious guidelines to follow around readability, encapsulation, etc., but these days I care more about system architecture than I do the code itself. Localized code is easy to change/refactor/clean up, the system itself is not.
> Ideally, when you have finished with each change, the system will have the structure it would have had if you had designed it from the start with that change in mind.
The cleanest code I write is for embedded systems without an OS, basically a sparkling gem of refactored goodness.
One thing I have noticed is that it's almost impossible to keep a codebase under control over time when business needs change and new people come in all the time.
I'm still trying to figure out how to apply this to my personal javascript codebases.
If I’ve given a speech, well there are several but the one relevant here is instead of trying to build the perfect product, building the best product we can build.
If you don’t follow that constraint you end up in Kernighan’s Law territory, and the wheels eventually come off.
Know your strengths. Build up or compliment your weaknesses, stop trying to Fake It Til You Make It when you’ve made it most of the way to where you’re going to get.
It's not exhaustive but it's a powerful general idea and I always like introducing developers to it for the first time.
I've found that the right level of abstraction is the one that saves time and effort and duplication now, for the features you're currently shipping. If you're thinking about hypothetical new features that aren't even on the roadmap then you've gone too far.
As an example, say your embedded program needs to load images, and your standard library only supports raw BMP files. BMP images are going to work great for the time being, since the art team can supply them that way. By all means, add abstraction method around the library method called "load_image" so that you don't have to refactor a million places to replace that library. It'll give you a great place to add error handling, logging, etc.
Beyond a single method to add a point of attack, don't go beyond that and spend the time to add a JPEG and PNG library, don't add support for high bit-depth images, or for grayscale images, or CMYK images. Don't add an abstraction that'll someday be able to load the Nth frame from a video file. Perhaps throw an exception for unusual input, but beyond that don't waste your time.
In my experience having simple, direct code makes it easier to refactor in the future when the need arises. Having too much abstraction or future proofing gets in the way because the future inevitably will bring different changes than what you were expecting.
Good abstractions evolve naturally out of good structure.
That doesn't apply to everyone and every project. When people leading the project are experienced developers understanding clean architecture the mess is just an oversight which is local and can easily be corrected.
It's a justifiable trade off for me, but I don't pretend that unit testing reduces complexity.
The very comment at the top of this sub-thread does not seem to limit itself to the subject of unit tests.
Same conceptual state gets represented in multiple variables or derived variables, and these must stay in sync. Very brittle
Oh, and have good testing in place to make sure you aren't breaking a required path that your IDE can't detect, obviously. No IDE in the world can detect "Oh, we still had one client on that old obsolete REST call and they are pissed"
What's a good name? I love the phrasing from _Elements of Clojure_ by Zachary Tellman [1]
> Names should be narrow and consistent. A *narrow* name clearly excludes things it cannot represent. A *consistent* name is easily understood by someone familiar with the surrounding code, the problem domain, and the broader [language] ecosystem.
At work, for any large feature, we usually go over naming pretty extensively, and aim to be consistent in documentation, code, and discussions, so everyone knows exactly what everyone is talking about.
But you also have to understand and internalize that it's OK to do a little bit of improvement each time. You don't have to go in, pick up a piece of code, sigh dramatically, and fix everything you can see about it. Just fix a bit. Turn some strings into enumerations or a custom type. Turn a recurring series of arguments into a single struct. Rename a deceptively-name parameter or function variable into something correct and meaningful. Add a test case for what you just did, or add a test case for something even related to what you just did that was not previously covered. Even just one of those is a good thing. Don't give in to the temptation to throw a 15th parameter on to a function and add another crappy if statement in to the pile of the god function. Don't fix the god function all at once, just take a bit back out of it.
If every interaction on the code base is net positive, even just a bit, over time it does slowly get nicer, and if you greenfield something with this attitude, it tends to stay pretty nice. Not necessarily pristine. Not necessarily nice in every last corner. But pretty nice. And if you do need to take out some technical debt, you'll have the metaphorical capital with which to do it; a non-trivial part of the reason why technical debt has such a bad rap is that it is taken out on code bases already bereft of technical capital, which means you're on the really bad part of the compounding costs curve to start with.
The example I keep coming back to is when I was a junior, one of the other juniors refactored the database handling code in one of our apps to use a class hierarchy. "AbstractDatabaseConnection" "DatabaseConnection" etc. And mind you this was on top of the java.sql abstractions already present.
I don't necessarily know what his end goal was, since the code still seemed pretty tightly coupled to how java and postgres handle connections and do SQL. One might theoretically now be able to create a testing dummy connection that responds to sql calls and returns pre-baked data. But the functions we had were already refactored to be pure functions, and the IO was just IO with no business logic.
Anyway, all it ended up doing was making it so I never touched the database code in that app ever again. Integration testing was handled by just hooking it up to a test db via cli args and auto-clicking the UI. And eventually when people started side-stepping it, I took the opportunity (years later) to just go back in and replace both it and all the side-stepped code with plain ole java.sql stuff that literally anyone with two thumbs and 6 months of java experience could understand.
So now, unless I have some really strong plan (usually backed up with a prototype I used to plan out the abstraction) for an abstraction model, I just write code, extracting things where the small-scale abstractions improve current readability, and wait for bigger patterns (and business needs) to emerge before trying to clamp down on things with big prescriptive abstraction models.
All globals should be configurable (most codebases I've seen have a ton of hidden globals).
All side effects should be isolated.
"Break any of these rules sooner than say anything outright barbarous."
We should stop seeing microservices as a technical problem / solution they are how to divide a "business domain" up into account the smallest constituent parts according to vat business view in the domain
The real solution is to recognize when new features are slower and messier to implement than they should be (because the evolving requirements have outgrown your original design), and periodically take the time to refactor to clean things up.
In a similar vein, when starting a new project form scratch, first do a quick and dirty prototype and then throw everything away and start anew. This way you know up-front what the challenges are.
About eight years ago I started a "proof-of-concept" that my manager asked for. I was using Lua for ease of development, and LPEG because it involved a ton of parsing. My intent was to get a handle on what was required and then do it C or C++. I found out a few months after the fact that my "proof-of-concept" was, in fact, in production and running. So much for my quick and dirty prototype. (And in retrospect, it hasn't turned out that bad---the code is way easier to deal with than our business logic in C/C++ because of Lua's coroutines make the event driven code look linear, and it's been fast enough).
Changing code that works, even if it’s a rock you ought to put down, is a risk.
Abstract interfacing techniques like base classes, abstract classes, and (my favorite) interfaces allow me to model interesting things.
Thinking about relationships between things in my systems versus categorizing things helps me avoid the "if you want to do something in OOP you must first define the universe" type problems.
DDD and conceptualizing how 'infrastructure' components interact with my main system is a nice guide for me.
Trying to write good tests is how I'm able to bounce around a few projects without having to read source code to reload context.
These are things that work for me. As I continue my practice I may find that I'm wrong or misinformed about some things. I should hope that I'll be able to incorporate a higher understanding as I gain more experience.
This is terrible advice!
Maximize your dependencies. Adopt as much external code as possible. Build what you can with it. Then, as you reach the limits of those dependencies, and you absolutely understand what needs to get done replace them as you need to.
The vast majority of what people write will be trashed and/or changed radically. You should adopt whatever tools are required to get things working minimally and then make decisions like this.
The libraries attempt to "dumb down" TCP, HTTP, etc and treat them as an abstraction that you don't need to know the details of. But it ends up biting people in the a$$ because networking isn't a perfect world where every request succeeds, terminates cleanly, or goes to the destination you expect. Whisking away all the complexity makes developers dumber as they eschew solving low-level problems with over-engineered high-level solutions that paper over the underlying issue, e.g. using mTLS to get around the fact that you're using DHCP to assign address space to nodes incorrectly, or making every API request a POST because the designer didn't understand HTTP caching, and so on. You get these endless problems that were solved decades ago because people keep trying to reinvent the wheel.
Doing more complex time and date work, then a good solid library for manipulating datetime variables will save your sanity. Need to right justify a string to set length then using a leftpad will leave you at the mercy of a random author on npm.
IMO there's a lot to be said for writing your own version that does 60% of what some library does, but 100% of what you need it to do.