Earning the privilege to work on unoriginal problems
landmines.substack.com
landmines.substack.com
They developed their own custom database, a shitty key value store that didn’t even have transactions because “Oracle doesn’t scale and it’s too expensive”. They couldn’t make reporting work so they copied the data hourly to Oracle anyway. (It turned out that Oracle does scale.)
They also invented their own RPC protocol, had their own threading library, used an obscure build system, an even more obscure SCM, etc…
Nothing was standard. It was all bespoke, webscale, and in-house.
When I started there as a junior they had me fix bugs in their code. Instead, I would simply delete thousands of lines of spaghetti and replace it with a call to a standard library. It was like spitting into a volcano to put out the fires of hell itself. There was no hope.
They burned $10M and made $250K revenue and collapsed.
Their competition used Oracle, ordinary tooling, and made tens of millions in profit.
And you know what? It worked amazingly well. I'm not just talking about scale, but also about how much any single developer seemed to be able to do. Their tools and APIs were tastefully designed, tailored to their own needs and fit together remarkably well. I remember that they were consistently building things with just one or a handful of people that would have taken multiple teams at other organizations I've seen.
Looking back, that was a remarkably productive engineering environment because of how much they were willing to do themselves. The key insight is that building something for yourself is qualitatively different from building something for external consumption or using something general-purpose that already exists. You naturally tailor the code and tools to what you need, which is much easier than building something for what you think other people might need—and, surprisingly often, easier than bending a different tool to your own ends.
But, perhaps more importantly, you develop a shared mental model of the system you're developing and how it fits in with everything else that you are doing. That shared mental model is, ultimately, worth far more than the code by itself... but it isn't legible to management, so it's hard to foster and preserve without a strong engineering culture backing it up.
The other lesson I took away from that experience (and others!) is that the high-level questions people focus on ("buy vs build", "monorepo vs polyrepo"... etc) are nowhere near determinative. Whether or not any given approach works out comes down far more to the little details and nuances of what you're doing and how you're doing it, as well as a healthy dollop of circumstances beyond your immediate control.
At the end of the day, if everything else was the same, but the startup you're talking about was using perfectly standard and popular tools, would you expect the outcome to be any different? I've seen far too many Enterprise-grade trash-fires to believe that. If the core dynamics that led to a mess don't change, you can have a "not invented here" mess or an "Enterprise Best Practices" mess or a "buy everything, build nothing" mess—but there's no simple strategic direction that won't get you some kind of mess.
> shared mental model of the system you're developing and how it fits in with everything else that you are doing. That shared mental model is, ultimately, worth far more than the code by itself... but it isn't legible to management, so it's hard to foster and preserve without a really strong engineering culture backing it up.
Sounds awesome, did they have engineering leaders with long tenure?
In a manner of speaking, yes.
They're big on FP and OCaml and do a lot of open source work around OCaml, and have done for years.
Speaks to consistent leadership.
My (100% outsider's) impression is that the focus, culture and teams at Goldman Sachs are totally different now thanks to market and regulatory changes, and that it is not a good place to work in a primarily technical sort of role any more.
I'd guess this is also an area where the Dunning–Kruger effect and misaligned developer incentives are involved. Inexperienced teams are more likely to take on reinventing the wheel. When you're starry eyed it can sound fun, plus its a great opportunity to do some resume building.
But as folks become more experienced, they can end up salty and disheartened. If you've been around the block a few times writing your own database might no longer sound so exciting. Plus life has become comfortable, you've mastered the tools and aren't really all that interested in taking risks so it's easier to just recommend IBM or Oracle.
Both scenarios aren't really all that great.
Maybe Jane Street and their ilk are just better at rewarding risk while simultaneously providing strong mentorship opportunities (I see the intern projects).
I imagine it's the quality of the team that matters, not the factors that go into whether or not to build bespoke in-house.
Jane Street is operating in a market where there are few off-the-shelf tools, the tools are your competitive edge, and hence it makes sense to design your own. Similar-but-different tools won't work, at all.
The place I worked at were doing the same thing everybody else was. They were making a pre-digital business digital, just like everybody else in the 2000 dot com boom. All the same off-the-shelf solutions applied. The secret sauce was the niche market, legal loopholes, and first-mover advantage.
The custom DB vs Oracle thing is a great example. The startup saved no money by developing a custom system, because they had to license Oracle anyway. They gained no scalability advantage, because Oracle scaled to their required level just fine. There's no edge there to be gained over the competition. Worse still, they wasted 3+ years developing their own system, which gave their competition a head start.
And they're in a market where a tiny competitive edge can reap large rewards.
They're rather the exception that proves the rule.
I'm sure all their competitors do the same for the same reasons. But for companies that view software primarily as a cost centre, I can see it going horribly wrong.
Trading firms tend to build a lot of their own tech, but Jane Street is special in that they build almost everything in house. Ocaml is practically their own language. No other competitor comes close to this level of owning the entire ecosystem, including programming tools.
This is the most important thing I've learned working for 15 years in tiny SV startups as well as the most successful company in the history of capitalism.
The main problem is you can't really teach what all those little details and nuances are (the tao that can be told...). You have to live it and experience it. Over time, you meet people who Get It (the founder of the aforementioned called them "A players"), and people who don't.
All you can do is focus on doing great work with the people who Get It, cutting those who clearly won't/don't, and investing mentorship in young people who seem like they're on a path to it.
> they were consistently building things with just one or a handful of people that would have taken multiple teams at other organizations I've seen.
> you develop a shared mental model of the system you're developing
How much of this was due using OCaml?
Despite some language-specific quirks, trying Elm has almost convinced me that dynamic typing is a mistake for any parts of a project which don't need to be used from a REPL context.
In particular, I've found that having an expressive type system makes it easier to develop and communicate this kind of mental model, as well as to maintain a close mapping between the concepts you use to understand the system and the code itself. I'm convinced that this is a powerful approach that can be incredibly effective and is hard to reproduce in "normal" languages, but it's hard to state this too confidently since it is entirely based on my own qualitative experience with programming.
Do you have any recommendations or warnings regarding general languages which reach in the opposite direction? Reason[1] and F#[2] are both examples: they attach pre-existing ecosystems and compile-for-$PLATFORM tools to OCaml-like typing.
OCaml itself is also intriguing. However, I'm concerned that suggesting it for non-personal projects will go over poorly. The "GPL" in its standard library's LGPL license may scare people despite both the linking exception and Jane Street's MIT alternative.
The worst companies I have ever worked for didn’t have the skills to do it.
So yet again it depends on the quality of your developers. Not how you do things.
Raised 6M, no customers, no revenue, no product, a buggy, non-functional MVP.
19yo founder was terrified to relinquish control to senior devs because he was scared they would implement something he wouldn’t understand, so the company sat there burning cash making tiny incremental improvements and flip flopping on priorities as the new-shiny whims of the founder changed every few days.
I have no idea how he raised 6M, I think it was a combination of lies and FOMO from latter investors.
I tried my best to guide him and give advice but his ego was out of control (but only backed by mediocre skill set) so I had to leave when I realised I couldn’t help anymore
Of course you gotta find the guys with more money than sense. They're not exactly thick on the ground but during a boom cycle there seem to be quite a few. (Often the winners of the last boom cycle, convinced that they were smart rather than lucky.)
my company raised a small round e years in. enough to keep us engineers comfortable but no where near 6 million. meanwhile my cofounders have years of experience in financial analysis and running a restaurant while I have 10 years of experience as an engineer before stepping in here as cto and first technical hire. We actually do have traction and getting new customers every day off a product that was self funded. I can't understand investing 6 million on a 19 year old with no experience with an ego fueled by insecurity. A startup needs to be run like a phalanx because there isn't room for redundancy and dead weight. And definitely can't afford to stick to a ego fueled "vision." I trust my cofounders to do their job and they in turn trust me. the "vision" should always be about serving the users and thats pretty easy to quantify if you are willing to humble yourself and engage with them.
the good news is we will have more favorable terms when we do a series A. The good thing about starting a startup in a recession is that you're more accountable to making your business actually be something people want.
People say this with poorly invested money all the time, but I must be missing something. You could invest it in an index fund, say the S&P 500, and annually yield over 10% on average [1].
[1] https://www.officialdata.org/us/stocks/s-p-500/2016?amount=1...
Until recently, this was less than inflation. A year ago the inflation rate of the US$ was more than 8%. So the net result when applying an interest rate of 5% was -3%. Less than in most of the period before 2021 when assuming an interest rate of 0%. I doubt that investors are not aware of this, arn't they?
Stuff like ci lets you ship safely.
On the flip side I have been the one working long nights and weekends reverse engineering code by engineers who prematurely built complexity into the system because they wanted to add a GraphQL api in addition to a rest API. All while in the pre-pmf days, with no value-add to the features that ultimately DID find pmf.
I do generally believe that cleaning-up after the pmf-hunting phase is itself a privilege that many startups do not get to experience, and should be treated as such. I understood the author as arguing that we shouldn't chase shiny things and should ruthlessly avoid complexity in favor of finding pmf. This philosophy is clearly illustrated in the devtools startup he is running. I thought there were some cool ideas there.
And I think that's because both things are true! This is one of the many hard parts about starting a brand new company, figuring out the right balance to strike on this. It's no surprise that companies mostly get it wrong, and in both directions.
There's no single right answer here. It depends on exactly what the company does, the exact path to product market fit, what growth looks like afterwards, and how lucky the guesses about all that stuff were.
There are no subpar products that people love. First of all, "quality" is relative. See chatgpt, it's wrong half the time, in a decade if someone were to release a chatbot of chat gpt quality we would say it's terrible. But today, it's the best we have.
The classic story is how airbnb and stripe launched without coding anything, everything was done manually.
Now launch an airbnb competitor today using the same strategy. Obviously comparing yourself to airbnb is dumb, because back then all software was terrible.
The actually successful modern companies of the past few years are openai, tiktok and figma. They all launched with complete products and are massively successful. That's what it takes today.
> I have seen companies die trying to perfect software no one is buying though, quite often.
Most companies die.
Until a product hits a minimum quality threshold, its useless. Which is basically what you stated later. Now most areas are mature enough that a copy pasted solution hits that minimum threshold (Eg, setting up a basic ecommerce site with ordering and payment). But for uncharted territory, that threshold is very real, hit it or die.
But my intuition is that failing to find product market fit (for whatever reason) kills companies earlier, whereas hitting product market fit with a subpar engineering foundation is more likely to slow companies down in later stages where they die more slowly or just underachieve.
I think Twitter is probably the best known example of that second pattern. It may be apocryphal that this was their problem, but either way, I think this is a real phenomenon in general.
The alternative to building your own DB isn't to put everything into a big json file.
The correct alternative is Postgres.
My point is just that it's very case specific, and you can easily guess wrong in either direction.
There’s a software law called Teslers law [1], which says that complexity cannot be eliminated, the most that can be done is to shift it from one part of the system to the next. You can make a similar law about risk. As a developer, if I ship a build system (complete a project successfully), or I build something with a new technology, I have a point to put on a resume and a successful accomplishment to talk about in my next interview. It may be that this is the wrong move for the business. It may be that they need a crappy spaghetti code abomination to achieve a product market fit in that time. If founders know what’s going on and they stipulate boring solutions the developer accepted risk. They’ve completely invested in the business and are completely tied to the decisions of their founders and leadership for their success.
In a perfect world this would start a conversation about compensation for accepting or balancing risk, but in my experience this absolutely never happens because it gets political. Leaders always win because they’re better politicians than developers. It’s easy for leaders to be willfully ignorant and dismiss these concepts as too detail oriented, and devolves into leaders saying “we’re nice people, trust us”. But the risk is there and someone is accepting it and developers respond to this by mitigating it on their end deep in the implementation details.
We have optimized at a local maximum, but are globally suboptimal. I do not believe we’ll ever achieve something more optimal because in business climates, everyone is trying to get something for very little effort. It takes much time to earn trust, but very little time to destroy it.
[1] https://en.m.wikipedia.org/wiki/Law_of_conservation_of_compl...
This can also help recruiting; that's why Uber had all those technical blogs about how super-complicated and full of Scala their solutions were, when Uber's business could be run on a PC under a desk.
Also why Palantir is called Palantir. They don't actually do anything cool or spy on people, but the name is intended to help get young tech people to work on boring government work without having to pay them more.
I feel the author is more criticizing the CTO who prematurely decides to rewrite the infra to handle scaling to millions of concurrent customers before they have those customers.
[0] https://www.dspace.com/en/inc/home/news/engineers-insights/b...
The trick is: learn CI/CD at your job. When you work on your startup it will be a trivial thing. I mean you probably will be using Vercel or similar where it takes 120s to setup and deploy from nothing. Most of that time is npm install on your machine.
What about other things? k8s for instance?
Simple: if it is second nature to you then you can use it. If you can’t fix 99% of the issues on the spot then don’t. Pick something else.
Making `docker build .` run on push is not a hard task and probably could be done in a few hours with cloud services. I spent a day to write a custom github workflow to run it on our runner, but that's probably not needed for most people.
After your OCI image is built, you can push it into the registry.
I guess that counts as CI.
Now E2E tests: I don't know anything simple yet. I guess I can hack together some script which will run docker compose and invoke that script with github push webhook. Probably would do it that way. Should work for not-so-complex projects.
If you have a couple engineers and you spend one engineer for one sprint to get things running automated for the next quarters push, that could be a good spend.
The MVP of CI/CD is a few hours wiring up Gitlab/GH Actions. The MVP of e2e tests is more variable but a moving skeleton can be hours if no UI, days if including browser testing.
However what makes sense for a first pass is very context-dependent.
I made another comment to this effect, but I think testing and CI (not necessarily CD though) have such quick payoff times in terms of productivity that you almost always want to set them up ASAP.
This was a hard-learned lesson for me, because I enjoy toying with build systems. If I can get a project into production from the start, I'm much more likely to come back to it and keep working on it.
> Develop locally with instant infrastructure, preview PRs in dedicated environments, and skip tedious Terraform with automatic infrastructure setup in your cloud.
If you thought this sounded like a sales pitch for bad engineering practices, it’s because it is.
I find it very funny that a company hawking their mutant CI “solution” talks so much about pmf when what they’re selling is pretty undifferentiated
> For engineers, this means taking shortcuts, making do with ready-made solutions, and focusing on speed of iteration over quality of iteration. You don’t get to spend time perfecting CI/CD pipelines or setting up the latest and greatest infrastructure-as-code versioning tool. You haven’t earned that privilege yet.
I agree, you shouldn't spend much time setting up the perfect CI/CD pipeline. That said, setting up a perfectly capable CI/CD pipeline in something like GitHub Actions took me literally less than a day, and maybe another day to work out some kinks.
I think it's important to focus on the author's primary point: The thing that is most critical is optimizing your speed of iteration. But too often I've seen this devolve into "we don't have time to build a simple deployment system!" or "we don't have time to write tests!" But the problem is that the lack of one-click deployment or any tests can become a huge drag on every single iteration in the future, and I've seen it absolutely sink companies.
Yes, things can definitely be over-engineered, and these days there are so many great, proven solutions for these common problems that it's a definite red flag if you build it yourself. But don't let that be an excuse for just generally shitty "we're a startup, we don't need to do XYZ" engineering practices.
And on the other end treat test performance as iteration speed so what once was a test suite you could run with vim-test on every little code change in seconds with a sqlite/redislite now requires docker, a 10 second startup time, and 5 minutes for a full run.
It's amazing the difference in productivity in projects where the test suite is fast.
It should take what, fifteen minutes on a familiar stack and up to a day or maybe two on something very unfamiliar that needs lots of web searches?
A dev can go an entire career never having needed to set up a CI/CD pipeline on their own.
Take GitHub Actions, for example.
That said, I do kind of hate this mentality; part of the reason that crap like Java 8 refuses to die is because companies always want to "established" tech, no matter how horrible of a fit it is.
A lot of these things aren’t binary yes/no decisions. They range from “don’t even do the thing at all” to “develop a custom solution completely in house”, sure, but with a huge spectrum of choices in between: buy off the shelf, get something basic but inflexible working, get something more advanced and flexible with more setup working, establish a process-to-be-automated that engineers perform manually.
Especially on the <easy setup, kinda shitty> to <hard setup, powerful/flexible> gradient you really need to consider your needs and options’ risk/reward. Some things like tests and CI have such quick payoff times you almost always want something that does a decent job of things. Moreover, when selling technical products half the time your competitive advantage is actually being able to efficiently iterate and improve on the product faster than everybody else because competitors are bogged down in tech debt and manual operations.
It does have, though, an important insight that I haven't seen talked about much: when you do some upfront optimization (e.g. for scalability), even if you end up needing it, most often your solution ends up not being fit for purpose.
This is true for startups, but it also applies to software in general. By the time your product is scaling, it will probably have gone through quite a few changes, and it's likely that what you predicted would be scalability bottlenecks are off the mark. Not to mention how damn hard is to predict what issues a product you haven't yet built will face on the first place.
Would love some talks covering where the adjacent areas are to the cloud offerings - are you thinking at AI inference startups or more dev tooling?
The oddest trend I have noticed (may be selection bias) is how many people build dev tools these days... to the point even non-devs are starting to talk about their no-code builder startups... it's getting crazy.
In ye olden days, we referred to this as “skill,” but I’m more sure we’re allowed to talk like that now :)
Been in operation for 3 years and scaling has come out though incremental improvements. no large scale rewrites needed.
1. out of the bod req/res time is fast
2. multiclustered websockets come out of the box with channels for subscribing to events
3. easy to spin up a "service" as a genserver
4. best ecosystem of tooling for building asynchronous and parrallel systems
5. liveview is THE standard for reactive frontend frameworks where you are updating in real time from serverside html. (no shade thrown on rails but hotwire is a joke by comparison) liveview processers are persistent for the live of teh user session and can subscribe to server side events so you can have a user make a request, push the task into the background and update the ui when its completed with very little code.
ok so type safety isn't at ocaml, rust or haskell level but its good enough that Its rarely the source of bugs assuming you perform even minimal testing.
It also faces impedance mismatch with a lot of cloud tooling--which is not a disqualifier by itself and BEAM might not even be wrong, but most Elixir-in-anger systems I've seen end up abandoning much of the benefit of BEAM clustering and running a bunch of horizontally scaled web applications because they've got containers to manage and it fits the overall get-it-out-the-door plan better. This makes things like "just add another genserver" not really map to reality all that well, though of course genservers are a good layer of abstraction and modularization on their own.
LiveView is great, though. I am hopeful for React server components and server actions to continue to mature; they're promising as it is.
you'll have to elaborate there. we have a pretty large codebase at this point that serves 3 different apps. granted our engineering staff is small but thats exactly what makes elxiir so great. It lets you build scalable infrastructure with a small team of humans
> most Elixir-in-anger systems I've seen end up abandoning much of the benefit of BEAM clustering and running a bunch of horizontally scaled web applications because they've got containers to manage and it fits the overall get-it-out-the-door plan better.
containers and clustering are not mutually exclusive. we run a elxiir cluster and take advantage of that with shared data and the ability to send mesages between each machine in pure elixir. At the same time, each vm is running inside a docker container managed with kubernetes. they work well together. we had one instance where a vm went down and kubernetes immediately brought it back up followed with it connecting to its peers.
In other words, I don't think there is any web tech that does this because it's all about what's between the ears of humans.
Engineers think a tech startup is mostly tech, but in reality it's almost no tech at first. In fact, the more tech you introduce early, the worse off you are.