Staying on the path to high performing teams (2018)
lethain.com
lethain.com
I encounter people all the time, that are inordinately proud of their infrastructure and process, and how that reduces team overhead, while allowing them to treat engineers like LEGO blocks. I worked for a Japanese company, and they actually make this work (sort of, and at a big price). It requires a fairly massive and inflexible structure (despite being labeled “agile,” it tends to be remarkably rigid).
There’s absolutely no substitute for a seasoned, experienced team that has worked together for years, and is secure in its cohesiveness. “Tribal Knowledge” is a dirty word in today’s business culture, but it is, hands down, the most efficient “team glue” in the world. It is also thousands of years old. No high-tech equipment required. Mammoth hunters worked as a team, and achieved big results. Military teams are as old as history, and that is all about commoditizing “tribal knowledge,” and investing in people, not just weapons and tactics.
When I read “add slack,” I was like “How does adding a gab tool improve innovation?”. Then I read that it meant adding downtime/flex goal stuff, and totally agree.
That said, the military enjoys a level of personnel lock in that few software orgs can dream of. I don't know many software engineer that would agree to a five-year contract with no option for early leaving.
Which is why it's important to practice good, humane management. Good engineers can punch their own tickets. Need lots of carrot, and very little stick.
I kept high-functioning, smart, experienced, C++ image processing engineers together for decades on my team, despite mediocre pay, and a pretty challenging environment (no, I won't go into detail).
This required that I connect with each and every member of my team as a human, often on an equal basis, as opposed to as a "superior/inferior" basis. It also meant that I needed to learn about the drivers for each person, give them as much family time as they needed, back off from micromanaging, and find ways to help them to resolve differences in ways that strengthened the team.
As I have stated before, when they finally rolled up my team, after 27 years, the person with the least tenure had a decade.
In about a month, we'll be getting together for a cookout at one of my former employees' home. We haven't worked together in over 3 years, but we still share a bond.
I think that the article is a bit of a "mashup." You can't really have a team "falling behind," but a team can certainly be "innovating."
I actually disagree with this statement. Teams can be falling behind for a variety of reasons and throwing more people at the problem can often cause further issues especially if the domain / technology / code base / system ..... Is complex.
Onboarding people into a team is an incredibly costly exercise, to do it while a team is already struggling will often create a worse problem. You need to be sure that the reason the team is struggling is purely an hour's in the day and not due to other factors.
Otherwise, this read like a documentary about my current company and product.
Ultimately, if you're falling behind, there's not enough capacity in the system, and so some things won't get done. Better to make that judgment explictly by setting priorities, rather than having it happen ad hoc.
There's only so much work that can be parallelized efficiently before people start stepping on each others toes.
Even if you are able to split things into independently workable chunks, the lack of domain knowledge in a new team member will still often result in a net negative for the first few weeks/months
This is taking about the best way to minimize growing pains, assuming you want to grow your org.
> To spread hiring equally across the teams in need, or to focus hiring on just one or two teams until their needs were fully staffed? That was the question.
It’s not claiming that adding people will make the current project ship faster. It’s claiming that focusing hiring people to underperforming teams will eventually buy them the slack they need to start fixing their issues and increase their velocity in the future, and that you get less total disruption by hiring many people into one team than by hiring one person onto many teams.
Adding more people may not be the only option, but pushing back demands or changing the wider organizational dysfunction is probably not even on the table.
I think the author’s framework is that most “other factors” would be fixed if the team had more slack. So increase the team’s head count, wait for them to gel, and then with the new slack they should be able to resolve the issues.
I agree that there could be other factors (eg a toxic employee, say a TL or manager that is holding everybody back, or an ill-defined mission/mandate). But this article is really just answering the question of where and in what order to add employees if you’re doing hiring, rather than giving you a reason to hire more people.
Essentially it boils down to asking Why as much as possible until you get to the original problem. From there the problem can be put in the backlog to be solved by the devs rather than the precooked solution.
Maybe there was a wrong decision made at the beginning of the project that needs to be fixed. Maybe the version 1.0 has too many features. Maybe the project could be split in two independent parts (which will still require hiring more people, but with less complex of internal communication). Maybe the basic infrastructure was not set up correctly (logging, unit tests, continuous integration) and there is no time to make it right because the visible parts of the application always get the top priority. Maybe the team members are working with a technology they never used before, and giving them a one-week training could increase their speed dramatically. Maybe there is some interpersonal conflict in the team, where things could be improved by removing a team member?
More importantly, why are you adding new items to the backlog so fast? Is your company doing "agile rituals, but with predetermined scope and deadlines"?
Let me guess: the answer "hire more people" is so popular, because it is a universal answer that does not require the manager to actually understand anything; and as a nice side effect, having more underlings increases the manager's importance.
Oh, and a system fix is to add process? That's about as counterintuitive as it gets.
Especially disingenuous: "A friend is six months into supporting a sixty person engineering group"...you mean managing. Say what you mean, and don't wrap it in nonsense like that.
If the team is treading water, stop incoming work until they start making progress.
If the team is repaying debt but not at the desired capacity, then add more people. In that case, the effort of training and integrating a new person is just another form of debt.
In all cases, if the team is overloaded in some way (debt rising, not delivering quality, missing deadlines) also add slack.
Of course. It's the only logical and effective way to deal with the problem.
> It leads to an organization where everyone is essentially “blocked”
There's a big difference from being "blocked", and what the article called "falling behind":
>> "A team is falling behind if each week their backlog is longer than the week before. Typically folks are working extremely hard but not making much progress, morale is low, and your users are vocally dissatisfied."
So, you deliver one or two features, or consume a backlog item, but more work comes in, so there's net negative progress (as measured by the list of pending work items). So, you stop the incoming work so that the existing work can proceed to completion, the team can then figure out how to add capacity without falling behind again.
If the team is incapable of doing the work for some reason, that's a completely different problem.
...where, if you can share?
If you want to add capacity, fine, but you can't do it to a team that's already overloaded. When would they have the time to select, train, and integrate new people into their flow, if they don't have any bandwidth? So you put "add capacity" to the top of the backlog, and clear work until it can be done.
Don't solve problems you don't have. Most companies think they need microservices, for example. It is an extremely expensive proposition. Debugging anything becomes a nightmare and developer ergonomics are shot to hell. There is a reason why we stayed away from distributed services in the past. What, you think I could not have stitched together 20 different Python services on my machine in 2009 and put my productivity into paralysis? Don't fool yourself about things like Docker making things easier. It doesn't help people reason about distributed systems.
Instagram? Monolith. 12 people when sold to Facebook.
WhatsApp? Monolith. 32 people when sold to Facebook.
StackOverflow. Monolith, with a lite SQL ORM for highly optimized queries.
Look at that - all those things we did in 2004, still work like a charm. But heyo, go ahead and spend a few months learning Kubernetes while the lean companies laugh at you in the rear-view mirror.
A team that grows larger quickly needs some way to manage the complexity within the codebase and the difficulty of knowing whom to talk to. In theory, it is possible to build a modular monolith with clear ownership boundaries. In practice, this requires strong technical leadership.
My current company has a code base which, if we want to avoid falling into technical bankruptcy, teams will need to spend 35 hours per week talking to other teams. The goal of well-defined interfaces is to reduce that to 2-5 hours per week.
Saying “microservices” is a concise yet oversimplified way to describe what is actually needed: well-defined modules with clearly-visible boundaries.
Not sure what you mean by 'core' but my impression was that he was more of an evangelist and a community manager.
It's astonishing how many teams want to use cool tech instead of create a great tool for users or how many teams think that technology will solve bad management.
I really like the idea of Rust I just worry about how things shake out In The Real World(tm).
I have no doubt Rust is here to stay for a while so I think it’s pretty solid play when appropriate.
Today’s desire to make everything event driven and “real-time” boggles the mind. Solutions are contorted to fit the Kafka operating model.
The inevitable result is blog posts about how “we solved our self inflicted problems that would have never existed if we chose a sane design for our real problem. “
Oh and never forget that no senior stake holder ever really checks their info/dashboards in real time. Never mind actually making decisions in real time based on it.
This was incredibly useful as it allowed teams to figure out what had happened in prod without needing to wait for the next day's batch jobs.
OTOH, most companies don't need something this complicated, and I'd imagine you could do the same with MySQL/Postgres if you're not at FAANG scale.
It goes for pretty much all technology, that you need to choose the right tool for the job. For some reason the marketing just gets most people all worked up and ready to fit problems into technology that was never designed for their problems.
I see this picture quite often. There is a small team that managed to create a great product. Then they get investmens, grow in size, hired a few starts that worked before in FAANG like companies and now trying to adopt microservices. The initial team spirit is gone, people who worked there from the leaving
So your metric for succes is acquisition size, not good software design?
So yes, agreed: 'good software' is absolutely subjective.
My point about 'good software' was about creating non-bloated, modular software - software that is well documented and easy for others to collaborate on, contribute to, or take over.
I think it’s not so useful to look at 30-person monoliths though, as you say. I’m sure there are 30-person startups that are jumping the gun and building microservice architectures? Don’t do that. First build the monolith.
But I think what you are getting at in the latter part, and I agree: the question is whether there are 100, or 1000 employee-company monoliths. I don’t know of any, I’d guess there are some examples, but I’d be comfortable wagering that they are much rarer the bigger you get.
https://instagram-engineering.com/static-analysis-at-scale-a...
Even if another company can do it...that does not mean that you can. Know thyself.
Jocko Willink can lead a team effectively on 5 hours of sleep a night. If I try to do that, I lose my grip on the fabric reality and the meanings of words.
Know thyself.
* Instagram barely had any features when it was acquired - the entire API was like 3-4 endpoints, and all the web version did was literally _show a single photo_
* Whatsapp was extremely brilliant in its implementation/choice of language (erlang) - they beat competition's headcount 10:1 for the same scale due to that, but hiring Erlang engineers is no easy feat
Facebook is also famous for its monolith approach, so I wouldn't necessarily put Google and Facebook on the same page when it comes to the CS Gods they worship (-: