A Staff Engineer's Guide to Inventing Work
sujithjay.com
sujithjay.com
It's precisely because of this framing and mentality that platform teams don't actually serve people well and are usually highly dysfunctional towers of people inventing work.
The fix for having no market is to act like the teams you serve could leave. This whole article lists signals, and none of them is that. Being captive does not mean the users or internal teams don't have other options and don't notice. Being product-led means caring about your users. A platform teams should be product-led, not engineering-led in that very narrow meaning. Otherwise you invent work as this article so wonderfully exposes.
Empathy for other developers, trying to have a mental model of what annoys them about work, and a general product centric mindset are the real requirements, but few people that decide how to staff the team k ow how to measure those skills. And besides, when people have them, they are often sent to talk to the actual customer, not internal customers.
My efforts were often focused on cutting stuff like that out and using the company platforms directly, which was hard because there were reasons people had middle layers. The company platforms weren't bad per se, but they were in-house and hard to find expertise on.
And where do product people get ideas for new features? If you're Microsoft you look at what successful competitors are doing. If you're anyone else you look at customer pain points, exactly as described in the article.
Engineers should understand their customers and their product (whether internal or external), should understand their metrics, and should be able to make intelligent, informed decisions. Sticking a non-technical 'product' (aka marketing) person into the mix is generally net-negative imo.
They're just saying a good platform team engineer cares about their users.
Which is fairly at-odds in spirit with the quoted idea of "there's no market to lose."
The original article does say "talk to your users" but it also de-emphasizes this by having it among an apparent laundry list of other signals: crashes, costs
Focus on cost without talking to your users - aka the people who care about the spend? You might spend a lot of time reducing a number that isn't very important right now. Focus on crashes but don't talk to your users? You might address some things with easy workarounds ("i hit retry") while ignoring much more painful toil. This is captured in the details for those sections, but not in the headlines.
There's also some interesting stuff in there in some of the bullets, like overloaded use-cases and partner-to-prototype, but again, that's just more specifics on how to talk to your users.
I think the article would be a lot more helpful to a lot more people if it was a "Guide to Talking to Your Users" and then framed each of those specifically as: user discovery, and how to talk about the given part.
Spot on, I've worked in multiple small, medium and large companies. And these internal platforms almost always turn into ivory towers of useless churn.
Even if being forced to use it internally, in a lot of cases the projects ended up buying the same solution (or a working version of it) from an outside vendor when it came to fulfill actual customers needs on the outside.
If that wasn't possible we built what we needed ourselves.
If you have internal customers, I still believe you need internal incentive alignment. Otherwise what the platform builds and what is actually used and needed drift apart.
And you need at least one real customer project to dogfood the platform initially. Because most platforms start out way too generic with pure architecture astronauts [0] at the helm.
If you don't have a project-0 to test those assumptions against, you don't need to start that platform.
[0] https://www.joelonsoftware.com/2001/04/21/dont-let-architect...
The Rule of Three in Refactoring applies to this scale, too. If you are building a "platform" for one customer project, that's a one-off and quite possibly YAGNI. If you are building a "platform" for two customer projects, that's likely coincidence and still possibly doesn't show company-wide need. If you can build for at least three customer projects, that is a pattern, that provides real guidance on exactly how generic things need to be and better ideas how to abstract it.
The more dogfooding projects the better. And yes, rule of three applies.
"almost never a product manager handing you roadmap"
Oh well wake me up from my dream vacation.
My nomination for this months "clarity of thought and ability to express it eloquently" medal.
Show me a platform team with a full product setup and I'll show you a platform team with over capacity.
Not to mention the average level of PMs will take a team like this down if you can even find one who can understand the type of work that the team is doing.
Suddenly came a mandate to try to integrate our systems together, the first goal was to use their authorization system (which involved, I kid you not, setting up 3 separate EKS clusters that talked to each other and if the main one goes down, all go down) that the core Platform team of the company was setting up.
Worse even, the plan was to use us as "guinea" pigs for their new systems before rolling out to the rest of the company because our team was smaller and therefor wouldn't be as impacted by problems on their side.
I was vehemently opposed to this, all our other experiences with their stuff was just a huge amount of pain. I remember clearly a meeting we had where I said: "Why would I use your stuff that is not well documented and that I don't understand when I can use open source stuff that is documented and that I can understand". Their only argument was that said they would have internal support for us.
I came up with this concept that I tried to push on to them, that their stuff shouldn't be this monolithic tower of babel monster, but instead should be a buffet where I can pick and choose what to use. Our requirements were very different from their stuff and I didn't want to run into problems because of stuff that had 0 benefit to us.
I ended up leaving that job mostly because of this.
"If our internal design system is barely half as well documented than Bootstrap then teams will still use Bootstrap. We should take a page from Bootstrap's documentation and include as much detail as we can, on everything."
"As a platform team, your product is this API my app calls. My beta environment should be pointing to your Production, never your beta environment. If that means you need to support better multi-tenant from the same 'app', then you need better multi-tenant. Your outage impacts my testing, which impacts my velocity, which impacts my deadlines. Your backwards incompatible updates need to be planned and scheduled and rolled out like a Production update, every time."
Product mentality is still very useful in platform development. If you can't sell your platform on its documentation and its stability and you must sell your platform on mandate and top-down control, you probably aren't doing as much to help your engineering culture as you think you are.
Things never got too bad for us, but they easily could've, and it'd probably be like your scenario where I quit. Was lucky enough that a couple of big layoff waves didn't impact us but did drastically reduce the amount of bogus technical mandates coming in.
I agree. When I was on the platform team (despite it not being called that at the time) our thinking was that feature developers were our customers. And our production was the tools they used. This meant not necessarily aligning production being only what was in production for customers of the company, but also in any service we provided to developers. This mentality had to be instilled into the team, and we had to remind people of this frequently, but it worked well.
The problem with every platform team is that, impact is hard to nail down. Does X using your platform Y to generate revenue mean your platform Y has indirect impact same as X? Most companies don't think so and end up destaffing their platform teams. Until, the destaffing ends up in much more inefficiency because everyone is inventing their own crooked wheel, spending time on the same capabilities and detracting from product development.
No they have to be product led.
> Our PMs invent capabilities
this is not a good reason for it to be not.
IME as a staff+ platform engineer*, it is dead obvious what increases user value. What is truly challenging is building a story around why anything should be worked on at all, when platform sits the furthest from customers.
I previously worked with a seasoned PM from FAANG who had low technical chops but somehow placed themself in the final decision maker seat for platform work.
They relentlessly blocked work that I described repeatedly in details: data quality was hurting, latency was dog shit, etc. But “bad data model? What do our customers care about our data model and pipelines that are confusing to maintain?”
It wasn’t until I showed some basic charts about our core database being oversubscribed and at risk of a more serious incident. At that point, we finally had the same understanding of the problem at the highest level and I was granted (lol) approval.
This admittedly took me almost a year to figure out. Others just trusted my judgement. The takeaway for me was really good tho. Even technical people probably don’t know wtf is going on in your domain, and metrics gets everyone on the same page, because numbers and charts are easy to understand.
* Platform is used too broadly so hard to say what the author exactly means by this.
I say this not to diminish your frustration but to appreciate that you tried to meet him at his level, and provide some context for what the other side can feel like.
That is probably the case in most companies, and as I said it’s not that hard to identify what is impacted by that dog shit code.
Having a spike/PRD/TRD process, really anything written by a human long form, can help build alignment. This is because you have ample time to ask questions and dive deeper, and it’s not personal it’s just part of the process.
I’m not sure whether the charts I created were themselves the convincing piece of evidence. They just happened to be where I found common ground easier and then maybe the rest just clicked
Those are just not to convince a manager though, it gives numbers to set KPIs, or at least some benchmark to evaluate the improvements, and at the end of the day solid reasons to justify promoting the team.
From both sides of the fence, fixing problems that only the engineers on the ground can properly value leads to everyone being miserable in the long run.
What were the charts that convinced them? Just high load % on your databases?
This is so strange. To my mind the only purpose companies hire engineers is to support business. The staff engineer should not need a project manager to tell what is interesting for business aspects - even though the goals are likely mostly technical.
I do realize this does not hold up always. But to me if you can't provide some reasoning for your work in business metrics you are participating in an academic exercise.
The OP is discussing tasks to do, you sound like you are discussing things that need to be done. The OP isn't saying there aren't things to be done, they are saying that no one is going to tell them what needs to be done, so they must discover and formalize it.
"Inventing Work" is a bit tongue-in-cheek.
Lots of hires, especially in places big enough to have dedicated teams for platform work, are for empire building, rather than to support business needs. In these situations, companies hire engineers because a higher-level manager needs to increase their headcount, in order to pad out their CV for their next move.
"Inventing work" is a strange phrase to use imho.
But for staff engineers, they need to to discover the scope themselves. That's where "inventing" coming from. Not saying inventing is a good word here, but requirement alone is not enough for staff level's work
I might send a junior to a well-meaning customer that already knows exactly what their requirements are to just document them and learn the process. Or I might require a principal engineer for the requirements engineering of: We need a new programming language. Let's figure out the requirements for it and what abstraction level is actually feasible for the target hardware platform.
Same for writing code: A junior might write code, a staff engineer might write code. Just likely on very different levels.
If an engineer comes to a manager and asks for budget for work they invented, I'd expect the answer to be something along the lines: "That is nice, but can we please focus on the stuff we are required to do?" That alone would be reason enough for be not to put ideas that way, but to come up with a requirement why it makes sense to do the work.
A senior(l5) is the tech lead on the project level. He is not in charge of handling the team's scope. As long as he handled his own projects well, he would usually pass the review cycle. The difference is crucial. At staff and above, the engineer's responsibility is vastly larger than a senior's.
I am not saying a senior can't do a staff's job - that's how he get promoted once he demonstrated that he is operating at a staff level, i.e "inventing" scope for his team. But a senior's scope is much smaller than a staff's.
This applies to almost all US medium to large tech companies
The author used an LLM. The first paragraph has two em dashes (turned into hyphens) and a typical overlong list. The second paragraph has an em dash the human author turned into a semicolon, and the third has one he turned into a colon. The fourth paragraph does the colon substitution several times.
The fifth paragraph is so obviously machine-generated I don't understand why people are even discussing this blog post at all:
This is the other cost - invisible to many, sometimes including the ones who handle the toil work. Every team has toil, and it is almost never prioritized. But you already knew this one, so I won’t spend more words to say that toil is an important signal that there is work waiting to be discovered.
If all you want to do is post prompt outputs on your blog, I think a) you should be honest about it and b) you have an obligation to write very interesting prompts.
Might even be more useful to just share those.
>At one end sit the signals that arrive pre-argued
Wash, rinse, repeat.
Does the author think "product customers" hand you a tidy list of requirements the stupid programmer automatons just have to translate into code? Obviously not. Customers also don't know what they want, or could want. Famously, they claim to want faster horses.
Then the author goes on to list "crash led discovery", as if this was so different from prioritizing bugs in prod. Or, If you have no idea, just improving efficiency. Or talking to customers, err, users.
All of this is true and none of this is any different from any of the other software, just translated into other lingo.
For example
> Thankfully, the signals that help us to invent work are already out there, and they arrive from four directions - from the systems, from the users, from your organization, and from the industry. What follows is a guide to reading each of them.
In a good “ideal” org the only source of requirements should be users, aka customers, aka real people, user facing products, teams etc.
There must be no room for speculations and “engineering” (aka over engineering), nor performance review driven development.
I would personally fire anyone “inventing” work. Even if it’s myself.
It IS hard to try to explain an alien concept to someone when you yourself in fact understand it.
I also don't think it's excusable to say it's hard for others to understand. The good SWEs I've worked with knew how to explain what they do to partner teams.
because hn (the culture) rewards (with fake points) being contrarian.
Personally, I found the article well written. Platform teams ultimately want the overall company to succeed. Their leverage - and mandate - is through helping the engineering org ship safer, faster, and more efficiently. The signal recommendations mentioned by the author are spot on from my experience, and I appreciated reading their perspective.
There are many other source of requirements that matter a lot. One heavily underestimated one is employee frustration. A frustrated employee is less productive and is at risk of leaving the org which incurs huge costs.
A few other ones: environmental impact, social impact, security, ethics, data safety/privacy, regulamentory compliance.
You might say that all these eventually end up as a requirement from users, for example "an employee leaving because he is frustrated means it takes longer to deliver features to a customer". But that is a very roundabout way of looking at it.
I’m suprised that so many orgs invest so much in teams that provide no product value, rather than letting the product teams learn for themselves
At my last company it took ~4 weeks to get a storage container because they were so overwhelmed with work and blocking precisely the whole company.
Well, right now, I'm pretty happy and have a personal project that I'm very eager to have done and polished and see the result of. Hopefully life keeps throwing more of them at me. It's been a constant issue for me though.
...but I will admit, it does sometimes feel inventive, and when I describe what I do, it does feel... inventive... in some sense. Particularly "no one in the org asked me to..."... instead it's almost always, identifying and getting ahead of needs, or, responding to overlooked friction/problems, etc...
Two paragraphs later, the topic is:
> Cost
I mean this is just incredibly lazy thinking and writing. There's always a cost line to follow, just because there's "not a revenue line" means absolutely nothing it just sounds pithy.
And if you cut your cost hard enough, you will put platform stability in danger, or hamstring your engineering org's ability to test and iterate, which will make you lose your market.
Again, incredibly lazy thinking and writing.
when systems engineers through intelligent restructuring of low level data structures save 1 TB of ram for their company without loss of productivity that is unrelated to revenue and market but still improving the bottom line.
if FB puts strings shorter than 64 bytes not on the heap but into the string control data via a union, that's in the same category of saving costs without adding a single new feature.