Development slowness in big and legacy applications
michaelscodingspot.com
michaelscodingspot.com
Big Project Complexities - has nothing to do with software business. It's an organization challenge which has been known way before anyone wrote anything remotely complex with software. See planning fallacy, optimism bias, read about how car and structural companies has done it. Look up "GM, Toyota and NUMMI" as an example.
Cross Team Dependencies. Again, this is an organizational challenge. Cognitive loads, poor communication structures based on industrial org designs and army-based leadership with control and command. Hierachies etc. Read Fred Brooks and Mel Conway on these for a more specific software approach.
Meetings, meetings, meetings. Of course there is a "going around it" solution for this. And it's called a manager. No manager should allow any team member to attend meetings if no thourough agenda is available, no actions for attendees, no proposed actions afterwards and someone who controls the meeting. Meetings will then turn into writing instead of verbal.
Also programmers: Who and why decided that we should be doing this thing this way?!
Or: Why is the other team doing things this way?!
It's much like requiring that any computation should either return a value or produce an effect. Without that, it's just waste.
only know to be useless after having had the meeting, and the decisions or discussion were either moot, already known, or completely expected.
So how do you know whether a meeting is required, before organizing and/or attending one?
* Does it have less than 10 people attending?
* It there an here an agenda?
* Is there more technical people than managers?
* Do you know why you've been asked to attend?
If you can answer "No" to any of those questions, you're going to have a bad time. Now you can have meetings where you answer yes to all of the above, and it's still useless, but the chances are lower.
Being a bystander in a big meeting, that is the worst. If it doesn't affect me and my input is not needed, why am I there?
A meeting with more than 10 people tends to waste more people's time than a meeting with less than 10.
One structural cause of that is the “nobody owns anything, everything is shared and employees are replaceable cogs” paradigm. Because then it’s suddenly a real problem that people do things a bit differently. You have to not only agree on what to do (the interface) but also the how (implementation). It would be the same problem in a band if you decided to rotate the instruments every time you jam.
I say this as a manager: who do you think is scheduling many (or even most) of these meetings?
This works until you have an outage, disaster, or embarrasment of some kind, at which point everyone is banned from doing that without approval.
This leads to a kind of "trauma driven development"; when you suggest a change, you see a flash of fear in the eyes of people around you. If you punish people for improvements often enough, you don't get any improvements.
Honestly, you shouldn't deploy into production without testing. Development and deployment needs to be tracked and scheduled. Because worst case, it has to be reversible. If you start deploying directly, without coordination with others (skip all the meetings, right?), you are begging for desaster.
Convincing everyone that a mistake deploying software is equivalent to a botched surgery, though, that's a good one! I've not run into that one before but I imagine you could get a lot of mileage out of that at the right company.
Others rely on whatever software is produced, co-workers, customers, users. I think those people deserve a good quality product. And that requires at least some rules and procedures. Unless, of course, it is a single dev thing. In that case you won't have middle management and performance reviews neither.
That said, if you're sending people to the moon, and are actually working on the system that sends them there rather than the website, feel free to have a more rigorous testing procedure. But if you aren't, consider maybe you don't need it.
Not every industry has to follow aerospace, or life science, standards. Every company owes its customers a good product. And with complexity, product, organisation, requirements, that requires coordination, communication and rules.
Ignoring that, cowboy style, not going to meetings, deploying to production directly, not understanding requirements and blaming all of that on management is simply unprofessional.
It makes me sad that this is even considered a serious possibility. It implies that as an industry we have so successfully convinced our customers to accept and pay for junk - as if software being junk is somehow normal or inevitable - that making an effort to competently build a good product doesn't generate enough competitive advantage to be a reliable winning strategy.
Fast forward, today everything can be patched and fixed OTA. It seems that possibility made it possible to ship pre-Beta stuff. That all those OTA updates are an additional revenue stream didn't help.
The other big factor IMHO is the amount of speculative money that has poured into the industry in recent times. That has allowed a lot of businesses to survive for longer than they naturally could despite shipping junk and that in turn has taught users/customers that expecting better software is unrealistic. Even most unicorn tech businesses (FAANG etc.) still produce a lot of junk and little that is any good - relative to the almost unimaginable scale of resources they have available - because usually they reached that unicorn status by having one or at most a few massive successes that they can defend and that then subsidise all the bad decisions and junk products just as effectively as VC funding rounds.
When it comes to stuff like that, I feel old!
Even mentioning the true cost would shut down most deals.
I mostly witness hurried scenarios in software contracting where each contractor tries to do the bare minimum to meet the definition of demoable. They then leave the project and mark it as another success.
I wouldn't describe a well-made product with limited scope as junk though. By junk I mean compromising quality and taking on excessive tech debt, not just starting small because you have to start somewhere.
Meeting attendence is important, because
- others need input or have relevant information to share
- activities have to be coordinated
- decisions have to be met and shared
Getting the balance right is trickey, just not showing up is simply not acceptable so.
- There’s a record of decisions and the thought process behind them
- When asked for input I can take time to consider a response and therefore give a higher quality reply
- I can refer back to conversations or search them when I inevitably forget some detail
- I don’t miss key information because I was out sick or had to take a bathroom break
- There’s a paper trail you can copy/paste when there’s a disagreement in understanding
- In meetings it can be frustrating to try to get a word in: there are often too many people or just one really chatty bastard who won’t stop
- It’s easier to scroll past chatter about weekend plans and sports than it is to sit through it in a meeting
The problem with meeting notes is they only reflect one person’s understanding, and often the person taking notes only has half an understanding anyway. Meeting notes are like JPEG compression set to the lowest possible quality.
I work with a team that spans +9h and +16h from me and it’s great because there are few meetings and lots of written communication. It’s very easy to search past communications to find old materials or decisions, and that’s not just hypothetical: I do it ALL the time.
We’ll use meetings, but for more tactical purposes: troubleshooting a specific problem, or doing a demo of some feature (although even then I would suggest demos are better as a recorded video).
Most objections are really something like “wah, I don’t want to read or write, it’s too haaaard.” Reading and writing were some of the earliest inventions in human civilization and everyone learns them starting in pre-school. It’s a low bar for a grown professional.
>> Alternatively, skip all the meetings anyway and then fire your middle manager. Or, if that fails, go work for a company that values your time.
If developers have to commit and wait for a build + test + deploy pipeline to see if their code works, that should be priority #1 to fix.
My experience is that it just isn't prioritized, and you have to do it early to enforce good engineering solutions to keep it this way. It's much easier to keep something locally runnable than try to set it up that way six months into a project.
If you can't run the stack locally, it makes it harder to set up performance and acceptance testing later. Running locally also makes long build times more noticeable and it's easier to address these issues earlier.
Being able to run locally looks low value at first, but if you do it, and sustain it, you get extremely large project-wide benefits later.
my experience as well and it blows my fucking mind every time I see it. It tells me developers have _no clue_ what it is they're doing that is helping or hindering them.
I think a lot of dev shops just genuinely don't realise it can be done.
I've never been in a startup right at the beginning, but I don't understand how a stack would ever be allowed to transition away from being locally-runnable.
We were fine until we started using Snowflake. The SQL doesn’t quite transpile to Postgres, and anyway we’re using some Python functions where the equivalent SQL would be nightmarishly complicated. I hate that we can’t iterate fast on that part of the stack.
Could we build something that works nearly identically, but locally? Yes, but we’d need to submit some patches to sqlglot, which nobody has time for.
Could we avoid using Snowflake altogether? Only with a lot more work. It solves a few real problems for us.
Current project I work on, everyone looked at me like I was made of cheese when I asked how I spin up the project locally. They only do unit and component tests locally, the rest is assumed to work when deployed.
Another killer is no contribution guide as a living doc for the team. "Code reviews" end up being a weird gate-keeping mechanism to exercise your authority over others by pointing out things they haven't done "correctly" when there is no shared consensus on what is correct and why.
So in a way it's "local" because we have private environments to build & test the code (and don't need to go through a build system) but it is still a relatively slow code-compile-test loop.
fun times...
even more fun when that project is a GUI app installed on lots of computers and now you have a new dependency to roll out to every system because it would be too hard for the dev to revert and fix it where the code did not need that dependency...
Even more fun when that dependency comes from licensed software that you only have a limited number of licenses for
You can see the effect in many large SaaS/tech products in the form of latency in the double digits to forever looping redirects of their service to Narnia and back.
You really shouldn't have to do more if the contracts are strong-enough.
If you really are that fragile you have a micro-monolith not micro-services.
If you have to run everything, what is micro? The individual deployable binaries (but their sum is now even bigger...)?!!
Startup time is the main productivity issue in our stack. We load lots of datasets into memory on startup, which can take minutes.
But even just making sure that “somebody hits the button” is a 1 button thing instead of a 20 button thing can be a major win! Totally valuable work
On the other side of the coin: having good observability into prod (talking mainly about APM stuff and tracing) gives devs way more confidence in the changes they’re making as well. Make it easy to deploy stuff, but also make it easy for systems to log things, and make it easy for developers to poke into those logs (my favorite flavor of this right now is honeycomb, but mainly cuz you can do perf analysis and answer questions from PMs about usage quite easily)
The nice thing about all of this if you’re a mid sized team is that you can totally outsource this with the right kind of person. Somebody who comes in, has the right to write some automation code (maybe it’s Zapier even!), and is not part of office infighting about process. The mandate being “make this pipeline smoother”. Probably among the easier problems to fix compared to deeper business or product problems. And it’s also super visible to other teams (“wow the bug is already fixed?”)
EDIT: one thing I’ve seen done that “never works” (I have not seen everything) is splitting up a repo into smaller repos. If you find yourself thinking “with smaller repos things could be smoother” it’s likely that what you want is not multiple repos but a build tool like Bazel or Buck, that will let you describe multiple projects in one repo, without causing inter-repo friction come review time. Integrating these tools can save you loads of money on CI and enforce decent test suite separation. The unfortunate flip side is that at least Bazel is very idiosyncratic, and most other tools are based on Bazel. I have heard good things about Buck 2 but have yet to see it in prod
At Google we use Blaze which is Bazel and it works very well, as you say.
At Meta (where I was before Google) we used Buck2 and it was the best. Much faster and more laptop-friendly than Blaze. It was like Christmas when they switched us over from Buck1.
Meta is very very good at “20,000 contributors to this repo”. Google is about 85% as good.
Getting work done (especially on mobile apps that take 7 hours to build from scratch) at that scale without these tools would make me jump off the building.
I think a lot of it boils down to strong engineering, which is not getting mauled into submission by product / business. Strong engineering will naturally gravitate to optimizing long processes.
The other big factor which has many consequences is the overall organization culture - are people encouraged to be helpful, reach out to other teams without organizing meetings in the calendar? Is there a blaming/finger pointing happening? Are many disputes largely about ego, dominance, politics?
The "Meetings" section might be controversial, but I agree as well. There's just no getting away from lots of communication/alignment in large organizations, and synchronous/low latency communication should be encouraged. Yes, that means disruptions, but being on the same page/unblocking others is usually more important. The amount of lines of code produced is not the largest cause of productivity loss in the large organizations.
I can't really say I found it all that controversial. If I wanted to be controversial I'd say that most recurring meetings and any meeting with more than 5 people are a waste of time.
I quite often ask someone who messages me out of the blue with a problem to book time in my calendar (which is usually at least 50% free). Because having a specific time to look at something is usually much better than me either dropping what I'm doing to think about something else or forgetting about their message and getting chased by them 3 days later.
There's no hard and fast rule, but if you can help them out in 10 minutes, then it's much healthier for the culture/organization to unblock them immediately (within let's say 30 minutes) than ask them to organize a meeting.
You're probably right but it's sad that this is regarded as controversial rather than self-evident. Maybe any meeting with more than 5 people being a waste of time is a stretch and most again would have been a more defensible claim? But the spirit is right. Good meetings happen when there is a good reason to have a meeting. Recurring meetings are often symptoms of other management problems that have better solutions available.
The real problem is often that people are getting blocked so easily in the first place. I'm sure I've seen every excuse imaginable for why organisations need a constant stream of pings via Slack or Teams or whatever they use. There is usually some underlying claim that everyone should respond to those pings quickly and accept the endless random interruptions that cripple their individual productivity because it's somehow better for the team as a whole. But if you make the effort to plan things and coordinate properly to begin with and you allow your people to concentrate and get on with their work then sudden, completely blocking problems should be unusual and your processes can be chosen accordingly.
That assumes something like this is even possible, won't be completely wrong immediately after the plan is finished, and won't take weeks / months to finish. This is what we used to do in waterfall - architects created a detailed specification which was obsolete / wrong / pipe dreams the moment they hit the Send/Publish button.
I have however seen many projects where management did understand their product and their market well enough to plan strategy more than 24 hours in advance. Their developers could then be reasonably confident about what they'd be working on for a while and act accordingly. If your organisation is functional on this basic level then getting the right people together early to gather information and consider alternatives and sketch out a design is useful for the same reasons that getting peer review when the code is done is useful or that having more than one person writing tests or doing QA is useful.
It's far more efficient to identify and correct problems early in software development. If you skip the up-front discussions and planning entirely then you miss several useful opportunities to do that. Typically in that scenario the problems are either identified and corrected during code review at the earliest or - worse - they are not identified and corrected at that stage either because the right people didn't have the right information to see the problem or because they did see it but it was judged too expensive to fix by then - and so the resulting bugs and tech debt make it through.
Yes! "Stack ranking" and inter-departmental competition and all these other dog-eat-dog things turn your employees against each other. People will waste unlimited amounts of company money to prevent another team improving their situation.
> One idea to encourage helpfulness is to include it in a company’s mission statement. Or better yet, add helpfulness to every employee’s performance review goals. If money depends on it, you can be sure people will make an effort.
Noooooo! You've gamified a metric! You know what happens when you gamify a metric? You destroy its value as a metric. And encourage people to cheat it, again. You've made money contingent on appearing helpful, not on actually being helpful. This is a great way to get interdepartmental meetings that achieve nothing but that everyone can cite on their "helpfulness score".
This is one of the major ones for me. It's a multiplier for everything else. local builds, CI builds, feedback cycles. If you get your build times down, everything else gets faster.
Migrating to Bazel was a godsend, while Bazel isn't perfect, it's one of the best build systems for large C++ (or mixed) code bases I've seen. buck2 looks promising, too, in this regard.
The single phase build is amazing as a concept.
I feel like the dev experience - without internal help for initial setup - isn't quite there yet, especially regarding error messages and documentation.
The way I like to handle these systems is to slowly transition things into “microservices”. I put in in quotes because it doesn’t necessarily have to be what many think of when they read the word “microservices”, and in truth they aren’t at all since they often build upon complex shared data models with all the “fun” that comes with that. Eventually I like to build things in to actual microservices, but it’s often a loooong if not impossible struggle to get there within most organisations which view IT as a cost center on par with HR but without the communication charms. Anyway, what I like to do is to separate complex business logic into their own “microservices” so that anything related to X lives in a single smallish services instead of a big complex system. One which is example is how we’ve got a “budget upload and computation”, a ”liquidity (not entirely sure that’s the correct English word)”, and, a “bank account” service all of which used to be part of this huge complex system, but are now separated completely in terms of the responsibility for their particular businesses logic.
Eventually this breaks complex systems into a ton of minor services, which comes with its own set of issues and is obviously against a lot of what you’re taught in academic computer science because often you’ll have similar code in multiple systems, but it’s sooo much more maintainable and it’s so much easier to move fast and implement what the particular teams in an enterprise organisation needs for their particular business logic. In my example, we’ve gone from five teams having to wait literal months for changes to being able to implement them within a week, and often the same day they are requested. We’ve also lowered our incident reports on problems from less than 60% to above 90% each month. I’m not sure why we call it lowered when higher is better, but whatever.
Just would be true just through survivorship bias, no? The systems that were already unmaintainable 30 years ago could not be maintained and either fell over or got replaced, while the maintainable systems from 30 years ago are the only ones left.
Occasionally someone suggests a major rewrite but that usually fails when it doesn't reach feature parity, nevermind meet the expectations of all the new features which were thrown at it to justify the expenditure.
Well put!
You can't just say it's self evident, because the people making these decisions haven't complied code in years.
Who is that parental figure? Why have we reached this state as an industry? Why do we have to beg for time and space just so we can do our job right?
To me, the problem is in today's management methodologies there are too many people loosely related to software who nevertheless participate in the process of software development.
Scrum master bothering my teams with some rituals that loosely relate to software but are actually tailor fit for a Japanese factory floor from the 80's, and have a weird cult-y smell? Out of the door with scrum.
Product owner wasting the time of my teams with childish poker card games and refinements and sprints and calling the shots on priority, considering they are the least technically qualified in the team? Why do we give the decision what to build to the least knowledgeable person? This is insane! Out of the door with this guy.
Agile coach (yes, that exists) just talking and talking about how this is not the right flavor of scrum? Out with this guy also.
Managers or sales attending the dev meetings trying to pressure them into delivering something faster? Throw them out and very sternly explain to them to stop with that, things will be done when they are done.
But wait, how do developers then know what the business wants, you ask? Well, I am glad you brought it up.
Engineers should decide what's the best way to build something awesome. They can do that by requesting feedback from sales or business representatives, talk to clients, view usage metrics, whatever. Whatever it is, it's their business and responsibility how to do it, what features to develop and how long each feature takes. They should also damn make sure the product is stable, free of bugs and reliable. Nobody else should have a vote in making engineering and build decisions. That's it. End of story. C łevel dudes better take a note of that, cause me and my teams, we are paid to build an awesome product, so leave that to us and don't waste our time trying to interfere.
If anyone who is not an engineer has objections, I don't care. I don't go to your sales meetings (those things are noisy and chaotic anyway) and teach you how to do your sales magic, right? I don't make you play poker cards trying to squeeze out of you how much time till you make a sale, right? And I damn right don't tell you what client to focus on and that each client has to be acquired in a 2 week period, or you somehow failed the sprint. And who can sprint for two weeks, no rest, and sprint again? That's just harmful bollocks. So why should I accept a non engineer having any say of how and what engineers are building?
I am at a fairly high executive position in a scale up, but that has been my approach and line of thinking for all my career. Remove anyone not engineer from having any say on what engineering does and you will build good, stable, reliable and sellable software. Even the guys who pay our salaries, they hired you not to execute their instructions like a trained monkey but to use your skills and expertise to do it yourself. So if you find that some management types are sniffing around you and your teams trying to make you make too many compromises, remind them that road leads to you using only 5% of your potential, and if they don't listen, find another place to work.
We can build good software if we focus on that and push away all constraints trying to make us do otherwise.
Also I'd love working with you (programmer with 22 years of experience, 43 y/o).
My question to you is: do you manage to implement the policy you are advocating for?
Thank you! I feel like it had to be said.
> Also I'd love working with you (programmer with 22 years of experience, 43 y/o).
Gladly! Message me to <my-nickname-> at protonmail dot com although I gotta warn you my hiring budgets are frozen atm.
> My question to you is: do you manage to implement the policy you are advocating for?
Yes! I got lucky enough to join a startup early-ish in my career and scale it up over the years. In my position (VP of Eng)I managed to have enough of influence on the the money people to be able to run the engineering org properly and without much interference. I even managed to pull off a 4 day work week for the developer teams with no reduction of salary. People've been happy, rested, our retention is very high for the industry (+4-5yr) and productivity (I hate that word) didn't suffer a single bit. Who knew? (Apart from 100s of studies).
Anyway, all that came with a cost of me not being very popular with the C-suite of the company, but that's life. You can't win all people's hearts, you just gotta try to do what you think is right.
Yes. Regularly and repeatedly.
> I gotta warn you my hiring budgets are frozen atm
That's quite fine, I just started a new job beginning of December 2023 and it's going to be a while until I start looking around (likely part-time / weekends but who knows, we might disagree on the promotion when the time comes in 3-6 months).
> all that came with a cost of me not being very popular with the C-suite of the company, but that's life. You can't win all people's hearts, you just gotta try to do what you think is right.
I wonder if "popular with the C-suite" is as good as many people make it out to be. If you achieve results then the fact that you're not singing everyone else's tune should not matter... but I've lived long enough to know that many people will view you as an enemy of some sort; a wildcard that must be tamed and molded and reshaped to be like everyone else. Because frak your personal approach that actually works, right?
Humans' social behavior is both savage and puzzling to me. Appreciation of personal style very rarely exists in work.
Anyhow, huge topic. I'd love to discuss it after I ping you over the email. :)