This is often the problem when hiring too fast. New employees don't have context or direction but generally people try to be productive, so they start inventing work.
Management team is new too, and not built up either direct or handle the bandwidth and they also are lacking the direction. Senior leadership and founders are likely busy hiring more folks and potentially trying to sort through the chaos described before.
IMO, pre-product market fit company shouldn't have 500 employees, maybe not even 10. These things don't become easier with more people, usually everything gets harder and slow. You also shouldn't have massive marketing or sales force before you actually have product to sell and have figured out your sales motion.
It seemed that Fast was just trying to shortcut their way to be the next Stripe by hiring people instead of actually working on the fundamentals like the product and business.
If you have a 500 person org that has operated for years you can probably add 10%-20% new people each year without it breaking too bad. But it can be very hard to build a complete function from 0 to 100 people within a year because there is no infra, culture or understanding how things work. The new people cannot be absorbed/aligned because there is no central gravity so people easily end up pulling different directions and spending their time on internal coordination vs actual useful output.
The parent mentioned 5 teams instead of 2. That's director or higher level decisioning. Which is where you frequently end up having someone far enough away from the actual work that needs to be done, starting early to come up with a roadmap. Given fixed deadline and fixed scope, you lean on your PM triangle knowledge and try to get as much headcount as possible to ensure success, because you don't have enough knowledge yet to really decide at what point adding headcount leads to a negative gain. It's usually less about empire building and more about trying to plan things early rather than growing in response to need.
I've been in this situation, as a manager, being told to hire three teams (two being teams of contractors), pushed back to say I should hire only one (the FTE team) as we didn't have enough work for them, was rebuffed and told I needed to hire (the roadmap said we had the work!), hired, and then had leadership panicking about the fact the teams were idle with nothing to do. After a couple months had enough work we could split it (unnecessarily finely) between them so they could all look active, which lasted about a month before the FTE team just started doing it all, as that was more efficient. I didn't stay there long.
This may be true about Amazon but it is not a given in general. Managers can operate in the grey area where they deliver just enough to justify more headcount "in order to deliver more" where the underlying motivation is increasing headcount. Not atypical in political environments like banks etc.
For your huge hires, 1/3 of the people were hired to be fired in the stack ranking, which of course is a huge morale hit.
The ones that aren't setup to fail lose 20-50% of productivity doing machiavellian backstabbing, defensively or offensively, for promotions other culturally ingrained/highly encouraged organizational sociopathy.
With 50% turnover per year (IT stats, not warehouse worker stats), it'll be hard to sustain development as well.
So it makes sense that Amazon would overhire.
Get out as soon as you can.
Heavy stack ranking basically means for a life event, a worker is very likely to get axed. Any health event, having a child, partner has hard times, or management reorg that puts you with a boss that doesn't like you, and the writing is on the wall.
Companies with a good long term focus should concentrate on acquiring and developing long term workers.
Amazon in particular structures its pay (payoffs come from bonuses and vesting 3-4 years in) in such a way that if one takes a perverse sociopathic view of how they manage workers, that management is incented to rug-pull workers three-four years in so the vesting doesn't happen.
Amazon is DEFINITELY sociopathically and perversely managed to do this.
Once that fundamental disdain of your work force becomes ingrained, then all other organizational abuses (stack ranking, hire for fire, backstabbing) is all fair game.
It starts with stack ranking and the fundamental management sociopathy. Which anyone that looks into accounts of Jeff Bezos becomes apparent of the source.
So why has Amazon been successful? Amazon's advantage in the overall corporate competition landscape was simply that they built an integrated business+IT strategy while every other company treated IT as a cost and annoyance.
But they treat their people like utter shit, and as Amazon transitions from an explosive growth company to a more sustained presence, and from AWS from explosive growth to "utility", the now-huge middle management and workforce devolves from stack ranking and sociopathic disdain into a cesspool where only the wicked survive.
https://en.wikipedia.org/wiki/Toxic_workplace
Unfortunately, Amazon continues to ride a growth momentum, so the toxic workplace will be viewed as a positive. Nothing will change until serious incidents occur.
Of course Amazon is huge, so there are workers who may be in less troubled areas. Just keep in mind that toxic areas of companies that derive from corporate practices will inevitably spill over to the "good" parts. Amazon is acquiring massive amounts of pure sociopaths in their management culture, and those sociopaths will "eat" other idealistic parts of the corporate management.
Corporate america is adapting. Better integration of IT and management is going to come to Amazon's competitors. Amazon's apathy to its product fraud and fake reviews, now stretching into a three year ongoing problem shows that it is faltering. Its lies during its AWS downtimes show a blame culture in action.
My company's strategy is get one extremely solid person across all our verticals (Android / iOS / Backend / Infra) to minimise communication overhead and maximise iteration speed (since we throw out half the code we write anyway).
I would be very hard pressed to hire more than 2-3 people in each role even if we were offered 10s of millions.
In the 90's there was a joke that they'd take the engineering headcount and multiply it by 10M (or something), subtract a multiple of management and there's your valuation.
but ... but ... that makes sense ...
I have found that having fewer good engineers is much better than lots of bad ones.
That approach is treated as heresy, in today's tech industry.
I kept engineers (really good ones) for decades. They had many "life problems" (like divorce, cancer, etc.) during that time, and I kept them on.
It's entirely possible. I have done it. I ran a team that was all "top-shelfers."
I feel as if I was a very “human” manager, and that seemed to make the difference.
My employees were loyal to me, not the company. Not an ideal situation, but it WFM.
Depending on key roles is something that every company in history has had to do. It’s a risk, but one that has been paying off for hundreds of years, for millions of companies.
There are ways to reduce “bus factor,” yet still maintain high Quality work. The corporation I worked for, managed this, but the price is extreme levels of overhead and rigidity. That has its own risks. I was treated as a “cowboy” in our company, and was considered to be borderline “reckless.” Most folks here, would have considered me to be destructively conservative and risk-averse. I got the job done, though. Our team consistently delivered high-Quality solutions to difficult problems, for decades.
The “easy answer” is to keep expectations and demands on labor low. Keep quality to a minimum. Rely on “black box” dependencies. Don’t use advanced development techniques. Squeeze your coders like lemons. Treat every employee the same as every other one, and control for the lowest common denominator.
That’s not how I worked.
It's OK. The industry is safe from me, and my radical views. It was made abundantly clear, years ago, that no one wants people like me in their company, so I have been forced to work with people that can't afford even crappy programmers.
Poor bastards are living the nightmare.
It has changed a bit, but a large part of it is due to HR policies around salary bands. You can have one person do 10x or 5x and another do 1/3x but their salaries are often max .5 factor apart. Ive been in situations where i'd rather reward (perhaps with a vest period) one person with 3x the salary and just skip all the inter-person communications bs, but HR makes it hard.
We were able to do this at hedge funds, and i'm sure small startups can do this, but as companies grow, HR will generally not allow this.
The "shiny toys" in this case are employees and the task of managing them.
As a (future?) lead/manager, having more people on your team bolsters your resume. A lot of people would benefit from the company hiring more too as it generates demand in other parts of the business - HR, IT support, etc (which in turn causes more hiring if you need more people to deal with the increased workload). As a higher-up in the company, it looks better on your resume if it was a "big" company with 100+ employees rather than a scrappy garage startup with 3 people, and likewise for investors.
Basically, to move development velocity faster we "load everything". Seems like a really bad idea on the surface. In practice, we only ran into an issue with a single customer that was literally 1000x as large as the other customers. It took about two weeks by 1 engineer to rework a select few calls that were egregiously slow and we were back at it.
----
In other words, even with some arguably inefficient technical decisions, I've rarely if ever seen performance issues at early stage startups.
Massive servers are laughably cheap these days. E.g. 16 cores, 128GB RAM is €112 from Hetzner.
In the meantime have you looked at leaseweb?
https://www.leaseweb.com/dedicated-servers#US
I'm planning on running nimbus web services[0] there and there are a few other smaller providers I have earmarked depending on how brave you are:
https://www.defendhosting.com/usa-unmanaged-servers/
https://cc.delimiter.com/cart/dedicated-servers/
[0]: https://nimbusws.com
If you can make the problem go away by spending more in servers always do it.
For startups, the life and death is product market fit and you only get there by iterating on the business as many times as possible, not making something more efficient.
Put engineers on business iterations.
Spending time on needless infra (scale you don't need, AI you don't need, infra you don't need, things could have been postgres as in this case, ...) both prevents building useful stuff ($-generating features/experiences/...) and increases operational cost (maintenance, debugging, ...). Both the lost revenue growth and the increased costs are compounding. Imagine a maniac steadily piling on a bit more technical debt every day and eat a bit more of the seed corn every day instead of growing it.
A good phrase for this is "playing house": http://www.paulgraham.com/before.html
More like, sounds like management was trying to get as much money as quickly as they could from their VCs before it all imploded. The engineers aren’t at fault that management didn’t have a real product or vision.
I mean, the engineers held up their side of the bargain. They were hired for webscale, the company got webscale…
Now, if this was a more strategic startup, that had tried to hire pragmatic engineers, that insisted on webscale, I would blame the engineers. However, in this case, I think it’s pretty clear that the engineers did exactly what they were hired to do.
Note that this basically works 99% of the time now and is part of a learning process that had the whole thing taking 2 - 5 hours every day for several weeks, down to 2 hours a day for several weeks after that, and now down to the nearly inescapable 20 - 40 minutes a day I have now.
Yeah, not doing that would be my own form of playing with nifty toys.
I've done something similar before and it saved countless hours wasted on configuring databases, writing queries, fighting impedance mismatches. Postgres brings a lot of work. With a file you just need a couple of lines to write your memory out to a file and a couple more to read it back. Simple and effective.
Until it stops scaling, but at that point you've established that your business is successful enough that it warrants the effort of doing something shiny. Postgres has a place in this world, but no need to put the cart before the horse.
It's reasonable to argue that shelling out the US$800 for a server with 128 gibibytes of RAM (or renting one from a cloud provider for US$150 a month) amounts to buying nifty toys for the engineering department. But if you're paying each engineer US$200k per year in salary and a similar amount in benefits, that's US$190 an hour, so each engineer's workday buys you almost four such servers. Your RAM has a minimum transaction time of about 100 ns, while an SSD is more like 10000 ns. If you have 100+ GiB of data to query, the alternative to buying nifty toys is paying the engineering department to do optimization work that wouldn't be necessary if your data was stored in memory that was 100x faster.
The rough equivalency here is that adding one more engineer to your team costs as much as adding 16 dedicated Xeon servers at Hetzner with 128 GiB of RAM each.
Sometimes the best answer is still a relational database, of course! It's often easier to write your queries in relational terms than by looping over hash tables, though probably not if you're using an ORM, and Postgres is delighted to cache all your data in RAM.
At some point, you may grow beyond being able to run your production environment on a single server, at which point splitting things into tiers is worthwhile. (In some cases you start at that point, for whatever reason, but a lot of web services can serve a million users on a single server. It sure sounds like Fast's service would have worked fine on a single server.)
Programmers love to add the latest tech to their resume then move on to a new position within ~18 months
If/when a developer builds deep knowledge of a problem domain, things like framework-of-the-month may become an unwelcome distraction.
Well the best way to reap the benefits of your experience is to run your own company. Not screwing yourself over with BS tech is a competitive market advantage. However you might need to hire people so you can’t pick something too old.
Working for other people it’ll mostly net you less on call headaches.
I've seen a few startups that seemed to overengineer things (or build custom solutions rather than use something off the shelf) in expectation of massive growth.
I also thought it might have been due to resume driven development and the need to keep engineers engaged so they didn't leave.
I've also seen successful startups with a monolithic rails or PHP application that ran on heroku with next to zero custom architecture components.
My experience has been that companies who will customize integrations for you have poorly designed products. By that, I mean, when we ask for a feature, and they bring on an engineer who does some requirements gathering then comes back with some invisible solution. That's a major redflag.
The success of an integration, in my experience, is directly proportional to the amount of the work I'm able to handle as a client.
I'm feeling a bit called out right now (not an owner but that's my experience, too)
I've seen this repeated on HN, but never saw the alleged flak.
They went from a JSON file, to etcd, to SQLite. etcd seems a little misplaced, but presumably it was already in their infrastructure and they thought they could save time leveraging it. The file-based approach seems appropriate for their particular use-case, though. It's not like it's a Rails app.
I’m still of the opinion that whatever you are storing, if your alternatives are JSON or Sqlite, etcd is a really strange/unconventional choice.
Tootally worth the time /s