Mistakes, any mistakes, are very hard to excuse when people, or other financial entities happen to lose money.
Comparing it with a site like reddit is, frankly, a bit disengineous.
Your customers are paying for speed and performance .
In a industry where users spend millions to be closer to the exchange if you are not provisioning for some multiple of your peak load you are going to get slaughtered in the market
The last blog post blamed overloaded DNS so it doesn't seem like a major architectural problem, which is why it's surprising that they're down again.
Nearly every write-action on Reddit appears to hit a queue. If your vote or comment doesn't go up for 30 seconds, that's not a problem.
If your trade takes 30 seconds, that could be a huge deal.
Somehow, in my mind, the responsibility of safely handling money of millions and redesigning the entire system don't really mix. Ditto for strict regulations and redesigning the entire system.
Please tell me you do not manage a team of software developers.
Rolling out fixes fast, even if they require intensive changes, is completely reasonable and expected in many industries though. They should have the talent and procedures to get it done. This isn't some startup web app, it's a multi-billion dollar broker managing millions in client funds. It's bordering on incompetent to have 3 outages in 2 weeks.
EDIT: What exactly is everyone disagreeing with?
This is a financial trading platform. Do you understand the risks of potentially introducing a different bug?
Changing a single line can introduce a different bug. Use proper QA and testing to catch as many as possible as with any development.
My emphasis is on getting things fixed quickly. They need to do whatever it takes to get systems online asap. Not sure what's so controversial about that.
I'm surprised by all the misinterpretation in this thread. Seems like it reflects the laid-back West Coast/SV attitude that isn't a good fit for high pressure time-sensitive work in other industries.
This is the part you don't understand. There is a difference between digging ten one-foot deep holes vs one ten-foot deep hole. People need time to plan how to coordinate and then get on the same page so that everyone can work at their own pace. That is the part that is not parallelizable and is the rate-determining step.
If this sounds unfamiliar or onerous then it's because you and others might have never experienced teams that do this. Robinhood is clearly lacking this experience and disaster planning.
RH wasn't prepared with any contingency. They should have a resolution for their users - even if they can't find or fix the original cause. That's the failure I'm talking about.
See the 2 other users in this thread that describe similar high-pressure situations.
Reality is, without knowing more about what’s causing this it’s impossible for either of us to say. If there is indeed some fundamental bottleneck that was previously not known, then I certainly won’t be surprised if it takes a while to sort out.
Now you can say they should’ve load tested, capacity planned etc etc. But we are where we are. Still can’t go back in time to turn this into a quickly fixable problem if it’s currently not.
Edit: also pretty disappointed that we don’t know more about the root cause. As an user I’d want to know what the issue was and what they are planning to do to about it to evaluate if I should trust them going forward.
So I'm not sure if I know that the state of the art for trading platforms is as rock solid as everyone is implying, and Robinhood seems to be way far off from whatever gold standards there are, see infinite leverage bug. So I don't think it's crazy that they move quickly to fix it.
I'd never be able to stomach the pressure, and I wouldn't wish it on others, but it doesn't seem crazy.
There are many ways to resolve these issues FAST. The easy way is to throw money at this. Go tomorrow (literally) and bypass all procurement controls and go to a mega big provider and scale this asap. Any financial services with a half decent IT has done the paper exercise to this scenario (at least for the purposes of BCP/DRP). The slow/better/mature way may be too slow, especially with the current market conditions.
The RH folks will definitely get a visit from SEC, their external auditor, and their external auditor will get a visit from SEC.(their auditor will be in the deepest of shits)(how come they failed to spot such a going concern issue?)(what the hell were they looking for on their audits?)(did they only send juniors over there?)
I feel sorry for the retail traders that got knocked down. I think that anyone locked in buying at 27-28k (US30) should wait 6 months to breakeven and after the US elections (irrespective of the winner) there will probably be a rally.
If this was Wells Fargo we'd see the pitchforks. But it's a SV company so all is good.
Not that I expect them to do any of this, just pointing out that 'shut everything down' is not necessarily the best approach.
IMO anyone who opened a position on RH in the last week should not be shocked that the same thing happened again in this high volume crapstorm.
Provision the resources and take the financial hit in costs.
If your architecture wasn't able to scale horizontally because it was poorly designed, heads should roll. This isn't some social network - these are financial platforms where literally the individual customer is financially dependent.
Heads. Should. Roll. Tell us the post-mortems, and then tell us who got the boot. Completely unacceptable.
A system like robinhood can have hundreds of moving parts, if 99 of them are horizontally scalable but 1 is not, eventually that 1 piece will become the bottle neck, and the fact it hadn't been made horizontally scalable yet is more likely to be a testament to how much work would have to go into doing so.
> Heads. Should. Roll. Tell us the post-mortems, and then tell us who got the boot. Completely unacceptable.
This is an unfortunate viewpoint. How quickly did we forget that robinhood is literally providing a service that no other company was able to do before. You want 100% uptime, you won't find it in an online service, let alone an online fee free service, you'll find it on the floor of the exchange.
ftfy
Haha, true. Call up you AWS account representative and you may find that certain service limits can’t be increased for love or money.
$0 commissions aren't a big deal either. If you just buy and hold, it makes no difference. If you trade actively then a real broker with better tech and order management is worth way more than the fees.