Software architecture pitfalls and how to avoid them
infoq.com
infoq.com
Love this thought! It's surprisingly common to meet engineers with 6-7 years of experience with incredibly bad habits that they picked up working essentially at a single company that they joined right out of college. Repeating the same questionable patterns for 6 years doesn't provide a lot of growth opportunity. In this context, FAANG's obsession with hiring new grads is a questionable practice, people get stuck.
Then they leave, and you find out the bits they rewrote were the carefully crafted and well documented protobuf APIs now replaced by ad hoc JSON where the parsing and spec are strewn across thirty places. The new projects you quickly realize make no sense whatsoever and don’t actually do the things they were supposed to do, but kind of look like it if you aren’t paying too close attention.
Now they’re a “senior engineer” at the next company.
very often development teams will value re-use and dry to the point that they'll conclude they need to stuff most things into a common library and share it amongst all their codebases.
Anyone who has seen this done long term has seen this accrete until everyone is afraid to change the common library for fear of breaking _everything_.
IOW, the risk profile of something like this grows over time. It starts out small but each time a project adds it as a dependency the risk grows.
and yes, I understand everyone is going to post about versioning and change management and all the myriad ways one can try and mitigate this.
the point here is that when using a common library like this you must be circumspect in what you put into it. The risk threshold for it being worth going into this common library is going to be much lower for someone who pushes for a common lib and then skips to the next job vs someone who was there over the next 3-5 years and personally experienced the pain and fear such approaches invoke longer term.
and to repeat myself since I know this is HN.
No one is saying you can never have common libraries like this, what's being said is that the long term risk is much more likely to be respected by someone who long term experience and therefore that person with long term experience is much more likely to be able to design a system that can be worked on productively over the long term.
a) Many people who have only ever worked at 1 or 2 companies have some really critical gaps in their technical breadth. Every time I've switched jobs, I've had multiple "I can't believe we didn't use this at my last job" moments.
b) I've seen many "lifers" in big companies who built some great big system at some point in the past and now do very little other than maintain this hairball. It's very rare that, for someone who has been in the same role for 10 years, their last 5 years were as productive as their first 5. Despite promos and all. You get older, the job gets easier, and you get lazier.
Somebody who never spends an appreciable amount of time in one role never has to deal with the long-term consequences of their decisions and misses out on that opportunity to learn and grow.
b) This seems like a problem you can screen for during the interview process rather than engaging in resume discrimination.
If one thing is a constant, it is that engineers will always come up with these silly little adages and cliches to discount the experience of their peers and continue telling themselves they are the smartest kid in the room.
These projects were the cornerstones of businesses sized from a couple dozen to a few hundred employees.
the point about being around for longer to see the effects of your mistakes assumes you're not a cog in the wheel.
> the point about being around for longer to see the effects of your mistakes assumes you're not a cog in the wheel.
Ideally—yes. In practice, decisions that should take QARs into account but don't are made by almost any IC level. I have seen systems designed by interns. You could say it's a company culture problem, and I would partially agree. On the other hand, tech companies have a tendency to lean into empowering ICs and so what ends up happening is that inexperienced engineers design systems that are only reviewed by overworked (and maybe not particularly experienced and/or motivated) senior ICs.
Startups are so chaotic and fast-paced that one usually only needs 1-2 years (if not less!) to see how earlier decisions pan out. Very frequently the stack and the codebase undergo monumental changes in that short period of time due to the changes in business requirements and scale.
Mega-corps are vast engineering efforts with hundreds if not thousands of daily contributions. While I could technically go back and try to evaluate my choices from 6-7 years ago, it would be fairly hard to decouple my individual contributions from the changes that happened afterwards (functional/non-functional feature requirements changed since then, the codebase is unrecognizable, etc). 6-7 years is just too long of a time frame for certain eng areas (web/native product is a primary example). Saying that, I can imagine that there are slower-paced areas where this time frame is more relevant, e.g. database engine development.
As a farmer, I get years of experience. I only see about one crop grow each year, so it really takes decades to start to see patterns in the conditions the crops come up in (e.g. years of drought, years with little sunshine, years that are cold, etc.) with little in my control to change that schedule. Good management decisions require having an understanding of all of those patterns, and as such, a year is significant boundary.
But in software, the most significant boundary is how fast you can work. The faster you can work, the more scenarios you can try. Time is not completely removed from the equation, but two developers with different performance characteristics can have wildly different experience levels after an equal amount of time has passed. In this case, a year is not a significant boundary.
I still think there ought to be formalized, international journeyman and apprenticeship programs for SWE, HWE, SRE, QA, product management, project management, and technical management independent of employer. The lack of generational knowledge and culture transfer leads to droves of novices learning bad habits.
"How to build perfect software in Django" is incredibly hard to write, but "15 common Django architectural mistakes" is a lot easier to write is extremely useful.
The unfortunate reality is that in many (most?) cases when inexperienced engineers face these challenges, it's too late to follow quick tips... their company's Django app has been set up 10 years prior, and now they have dozens upon dozens of layers of abstraction in the codebase.
Teaching how to think about architecture, how to evaluate options and how to make decisions is a more reliable and applicable skill.
The way I wish software were created is more like physical infrastructure. There's still huge problems with construction, to be sure. But it lasts longer and is more likely to succeed when completed. There's all kinds of requirements, analysis, and inspection to ensure it works correctly. The people putting it together don't need to be very skilled; they're working with off-the-shelf commodity parts, manufactured to a minimum specification, with specific dimensions and attributes, which loosely couple in many configurations with identical parts from different vendors all over the world. Combining them in specific ways has quantifiable, predetermined results. And you know that for the parts that require being designed correctly, the people designing them had very specific minimum qualifications that take years to attain.
Companies today, whenever they want to build a software product, think they need to build an entire factory first. But companies making physical products wouldn't do that, because building a factory requires factory-building skill, that has nothing to do with the widget they want to make. Instead they would find a factory and hire them to build their widget. Software may be "modern", but its production is antiquated.
This is not how it plays out though, you get burned from both sides. Inevitably some non-technical leadership stakeholders ask very fair questions about "why can't we just...", and they're not wrong about what's possible, it's just different from what came before and it's not possible to rebuild entire digital systems as quickly as the good ideas come.
On the implementation side, details matter and abstractions leak, seemingly small requirement changes undermine assumptions behind major architectural decisions. The idea that you can have a handful of seasonsed experts guiding an army of low-skill builders just doesn't work out with the same economics of physical construction. The details matter too much, and the implications of pure logic are too diverse to be covered by the equivalent of physical building codes.
I don't believe that. When you create a coffee shop, you build it to have a specific, repeatable, singular user experience. Some may like it more than others, but if the experience is really good, people will keep coming. You don't keep changing your coffee shop week to week for years on end. The maintenance is in things like cleaning the floor, not changing the shape of the coffee counter, tables and chairs.
> seemingly small requirement changes undermine assumptions behind major architectural decisions
That's just poor design. Commercial buildings are built large and open so that the business renting them has the flexibility to change things around in the space without calling a general contractor to rebuild the walls or raise the ceilings. If you build your software with tiny rooms and load-bearing walls in the interior, yeah, you're gonna need to rearchitect to get more tables in there.
The fatal flaw in these software systems is they're effectively a single business, who's hiring a general contractor, to build an entirely new building, with very specific use cases in mind. When their business finally grows, or they just want more natural light in there, they rebuild the building. It's just bad business sense.
In the very worst case we should be renting out a new building, not building or rebuilding one. Only the largest businesses should consider constructing new buildings, and even then, they should be building it to last for a decade or more. (And this isn't even about real estate, it's about their business requirements, wasting money on construction, and the risks of construction delays and failures)
Software is not like this as there is no limit to the scope that can be built in to a single app. You don't have the constraints of physical space, materials and human use cases and locality. The "load-bearing" aspects of software are not just physical infrastructure with CPU/memory limits, but also fundamental choices in the data model, some things more than others, and what starts out as a casual choice can become load bearing if further construction builds on that assumption. Buildings don't have legacy data.
> When you create a coffee shop, you build it to have a specific, repeatable, singular user experience.
I can't tell if you're suggesting software should be built this way, but after 25 years in the consumer space in multiple startups, scaleups, and largish companies, I have never seen any software company succeed based on rigidly adhering to singular user experience in this way. Even when you can describe the experience simply, the details and tradeoffs are enormous, and companies that win are the ones that are able to iterate quickly without letting the quality degrade to the point that the wheels come off. Facebook vs Friendster vs MySpace is a pretty good example. In my own founding experience, I was able to bootstrap a streaming service to 6-figure subscribers over 8 years leaving behind a graveyard of failed competitors because they couldn't thread the needle between good UX, good engineering, and the right content investment.
I'm sure it's different in well established and standardized corners of B2B and industrial software where you are not subject to the fickle nature of consumers, but in B2C you are default-dead if you insist on some kind of fixed requirement and long-term vision that doesn't change.
Isn't this in some ephemeral way exactly what most of us do? Sure, there are people in the world doing real Computer Science(TM) but most of us are assembling apps from existing parts. Sure there's glue and there are domain specific bits, but there's a lot under the hood in both paid and open source forms aren't significantly different from going to a parts bin and pulling the right SKUs.
Usually, it is better to branch a separate "new" team in another space, functionally deprecate the old project one feature at a time, and jettison the previous team including the manager after 6 months of uptime.
Trying to untangle a mess often takes 7 times longer than simply re-building a better version with well defined use-cases. =)
https://www.youtube.com/watch?v=6D9vAItORgE
=)
this is so true, but sometimes the incompetence is so bad it's not feasible to build a competent team (very very difficult at least).
I'm actually in the middle of this now and I have to tell you, I've been doing this for 25+ years and I'm absolutely flabbergasted that people with this level of incompetence are gainfully employed in this industry. It's so bad we have contractor teams running circles around them and by circles I mean 2-3x their velocity with code designs that are as good or better.
I seriously consider becoming a Plumber everyday.
Had to develop esoteric personal projects to find the fun. =)
Nope! Everyone can and does make mistakes. A good architect should accept and learn from suggestions and ultimately better technical decisions put forward by members of the team, whether it’s the senior with 70 years experience or the 1-month experience, bright junior with a good idea.
I was in a team where the architect chose to use React without TypeScript, and a nasty .NET 4/React mix for the front end. This was 2-3 years ago. I suppose it’s what he knew and was comfortable with? This combination caused no end of issues and an unnecessarily difficult development process. Unfortunately I wasn’t around at the time, but I’d have spoken up, and probably been listened to, which is what i’d expect from any good team, regardless of your position. Better is better, doesn’t matter whose mouth it comes out of.
A lot of the architects I’ve worked with have been rather snobby and quick to dismiss less experienced team members in my experience.
It is a nice idea, but totally impractical.
Um. Okay. So, just so I am clear, the advice here can be rephrased as: "Save a matter of life or death, don't have hobbies"?
And I am afraid I haven't broken your streak. I know what Django is and what it usually entails.
I'm not sure what "monolith" adds, though. Simply saying "Django" would communicate the exact same thing, no? Is there a "Django non-monolith" that it could be confused with?
Per your definition above, it seems there is no software that is not a monolith.
Per fallingknife's definition, that's still a monolith. No different than a Django application coordinating with MySQL. If it is that you think 25 > 1 is somehow significant, throw in Stripe, GPT, etc. in addition to MySQL. He would still call that a monolith. It's all the same.
That would not be a monolith by my definition as I say that coordination removes software from being a monolith, but we've long moved past my definition. We're only here now to understand where fallingknife's definition allows for there to be anything other than monoliths.
eg, I have worked with many so called 'developers'. They shipped lots of code, but most of it was solving the wrong problem, didn't fit any business requirement, added unnecessary complexity, had to be replaced almost immediately, etc. etc.
shipping software means compromise, most of these points are basically "don't compromise on X".
You surely wouldn't argue that companies should compromise on developer machine specs, for example? Too many organisations cheap out somewhere on the development process in the name of "efficiency", and this article is arguing (and I agree) that that's short-sighted and will cause more trouble than it saves in the long run.
I most obviously did not say, or mean, every piece of software shipped must have every single part of it compromised in some manner.
> shipping software means compromise, most of these points are basically "don't compromise on X".
as, you disagree with a lot of the article because it's telling you not to compromise on some things and compromises are necessary for software development. That only logically works if you're also saying that compromises are necessary everywhere, doesn't it? Otherwise what's to disagree on? If you're saying compromise may be required on any of items 1 through 10, and I'm telling you not to compromise on 1-5, well then 6-10 are your zones of flexibility.
If your stance is "always do X" then you can replace yourself with a post-it note. Just write "always do X" on the post-it note, then when you have a decision point, refer to it. You're done.
Actual engineering is about weaving through constraints and goals and that often implies compromises. Someone who can do that well cannot be replaced by a post-it note.
And guess what "always compromise" can also be placed on a post-it note so very obviously that is not what I'm saying.
don't be replace-able by a post-it note.