Robinhood’s third outage may point to deeper problems in its tech stack
techcrunch.com
techcrunch.com
And as an engineer, I used to point how Twitter's fail-whale was killing Twitter just before it got popular.
After almost a decade, now I get what he meant. And I'm more interested in jumping into a company with technical problems like these, with short time horizons rather than those with capital restrictions or chasing PMF in early stages.
I remember they were one of the first to really adopt Scala. They still use it quite successfully.
A rewrite may give you a temporary relief if the runtime of your new language is say 10% faster than your current one. This is not a long term solution though, and for many startups in hyper growth state, this will just give you a few more weeks of runway until you run into the next roadblock.
Your solution to scalability issues should almost always be system redesign, not component rewrite.
You're right, they probably are having issues with some lasagna architecture right now.
Yes, something with Postgres and PHP is probably ok for a lot of usecases, but a) Robin Hood is a pretty big target, and might exceed the simple limits of that, and b) there are plenty of cases where Kubernetes (or any other big orchestration framework like Mesos or Nomad or Docker Swarm) will simplify the codebase.
I'll agree that engineers are sometimes a bit too eager to just jump on the new HN tech a bit too early, but it's not like that tech isn't useful, and it's not like it doesn't serve some kind of purpose.
You don't know where the best place to fragment your monolith is until you've got to the point where there monolith is struggling to keep up, because then you know where the slow bits are and how to scale them.
And there's no need to split the whole thing. You can just split off the slow bits.
Starting with Kubernetes from the get-go seems insane to me.
Just one Nvidia Jetson Nano wouldn't be enough to handle all my server needs (movie streaming, video transcoding, random odds and ends of projects I have), and using a framework like Docker Swarm or Kubernetes makes this relatively easy.
Is having the ability to linearly scale my home server overkill? Maybe, but at the same time it certainly felt necessary for the job, and carries the nice advantage that I don't have to rearchitect everything later while trying to shoehorn in the "old" way of doing stuff for compatibility sake.
> You don't know where the best place to fragment your monolith is until you've got to the point where there monolith is struggling to keep up, because then you know where the slow bits are and how to scale them.
Sure, it's not a silver bullet, and I certainly wouldn't claim it as such. That said, there are reasonable places to expect bottlenecks that can benefit from an architecture; things hitting a spinning disk or making an external HTTP call tend to be slow, so it is often better to model them asynchronously and buffer through a message queue; even if you don't have concrete numbers on your side it's not an unreasonable assumption to make.
https://robinhood.engineering/building-an-application-deploy...
Eventually you will reach a point where you will need clever and fast code.
Dont get me wrong I'm not a fan of over engineering things.
But those two statements are not mutually exclusive. You can build scalable distributed systems with postgres as a backend, heck if I was building one that would be a core part of my stack...
I can't imagine our DBs holding up to millions of users or need to process millions of ticker symbols or trades.
What?? How could your team live with that?
https://robinhood.engineering/building-an-application-deploy...
I think everyone did similar templating for similar applications. For example Golang apps which have logging/metrics already built-in instead of needing sidecars.
> One notable project is to provide an intuitive, user-friendly frontend for executing GitOps workflows on Applications and Component manifests.
Is this supposed to automatically update manifests, deploy them to staging namespace/automatic tests, enable automatic canary deployments?
> We also hope to go even further and create a one-touch infrastructure-provisioning interface to abstract away manifests altogether, and place application-centric abstractions even more front-and-center for application developers.
How does that look like? Archetype-Specific web interfaces/manifests?
I'm kinda sad that this isn't available to the general public since I'd love to avoid constant repetition/writing my _limited_ own version.
YMMV but I think most brokerages would be happy to help if you ask.
If I were just starting out I probably wouldn't select Robinhood given the recent issues, but I don't see the need to change to a different platform at this time.
I also assumed that the recent outage was due to not accounting for yesterday’s circuit breaker. I imagined there’s a Robinhood product manager with a JIRA ticket who’s complaining about their feature getting depriortized every month.