Usually to aim for a significant promotion.
“Designed and built homegrown system to save $Xm! Give me promo, bro?”
Just so happened to ignore that it took X+Y additional to build. Also it will probably be going to the G graveyard in a few years.
Usually to aim for a significant promotion.
“Designed and built homegrown system to save $Xm! Give me promo, bro?”
Just so happened to ignore that it took X+Y additional to build. Also it will probably be going to the G graveyard in a few years.
If I were doing this, I would be looking at data warehousing systems. 1.7PB of, say, Parquet files in S3 is not terribly expensive. 1.7PB of Parquet files in on-prem or collocated object storage, even replicated a zillion times, is quite cheap. And quite a few companies and open-source projects are currently competing aggressively to provide awesome tools for querying that data.
The hot data would fit on basically anything — the choice should be about robustness and barely even consider cost per TiB. Datomic got written up recently and seems credible for this type of application. FoundationDB is bulletproof. Postgres could probably handle it without breaking a sweat, although active/active replication isn’t free. Heck, writes straight to a warehouse with a cache in front to help with reads seems credible — Uber rides rarely go for longer than a couple hours, and back-of-the-envelope math suggests that the total data rate is maybe 50GB/hour. An entire day of data for an entire country would fit on a single very ordinary commodity server, and the live data for the entire world would fit on one mildly beefy server. The indexes involved sound straightforward.
Old data could probably live at lower cost in a data warehouse, but then developers would have multiple systems and namespaces to deal with in order to query on transactions.
This seems like an odd requirement to me. What are the use cases beyond a few minutes past the end of the ride?
- Reconciliation of Uber's finances. This definitely needs to be correct and consistent and maybe even auditable, but it's also done at the end of any given billing period. A daily roll-up would be sufficient.
- Showing each driver and each rider their own history. This ought to be fully correct (especially for drivers), but it's a massive double-edged sword. Uber's app, for example, appears to be entirely missing the "please forget my ride history" button. And this isn't surprising: Uber apparently stores it in a ledger, fully indexed, complete since at least 2017, that is cryptographically immutable. Is this actually a good thing? What happens when a privacy regulator tells Uber to expunge, anonymize, or pseudonymize old data?
- Handling chargebacks? Having a high quality record going back 180-365 days seems useful. But the cost to Uber of occasionally losing a record is very low and is barely more than proportional to the probability of loss of a given record.
So I don't get it. If I were running an operation like this, I would prefer not to have a fully immutable ledger full of personally identifiable location, usage, and transactional data.
Exactly what he the OP said, authorization. Card networks cap latency of authorization requests (the time the merchant submits a request, to when it makes its make through the merchant processor, networks, issuer processor and back.)
What are you basing this on, banks have less real time constraints when it comes to certain data but thats rare. You will almost never see a bank, processor, etc. archive their system of record data.
Yes, which they already had as DynamoDB only held 8 weeks of data. Presumably all payments to drivers were calculated monthly and turned into invoices. Any corrections would end up as adjustments on future invoices. Seems pretty normal.
Promotion-driven development. I suppose better than blog post driven development, but marginally so.
Someone once called Airbus planes "Sun servers with wings". All businesses today have an IT department and those build around the idea of real-time scheduling done by computers are compute/storage businesses first. That's why Uber was able to launch Uber Eats or e-bikes. I am not surprised they have a bunch of busy devs building cool stuff. Not all of it is ego-driven. I liked their logger for Go last time I used it.
The architecture where all the magic lives in the database layer and the services treat the DB like this magical thing that synchronizes across multiple DCs for you, and takes care of every complicated part of engineering a distributed system is a "now you have two problems" situation. It's a massive complicated system that will fail in novel ways, and you don't have a large community of other users that you can fall back on for expertise.
That is one of the most important points of consideration when I am being asked to asses new tech. If the project is not actively supported by the community, I don't recommend it. I just can't expose the client to risk just because there's a new shiny thing.
I find DB development quite interesting, and I find Uber's core product quite not interesting (from an engineering perspective).
So as an outsider without any financial stake in the company, please keep writing about databases!
> Uber is famous for NIH syndrome
If you're truly paying engineers, project managers, etc $500k a head, it dramatically undermines the financial cost savings.
It very well might be the case that "We spent $25m of engineering resources to save $6m annually".
[0] https://www.levels.fyi/companies/uber/salaries/software-engi...
[1] https://www.levels.fyi/companies/uber/salaries/software-engi...
I made a config change to our AWS instances and projected approximately $10MM/year in AWS costs savings (pre-savings).
My boss asked me "Who told you to do this? We need to focus on $project instead". I found another team and transferred out. 3 months later there was a big fire drill about AWS costs and they took my 1-pager and executed it. Didn't get any credit in the shipped email nor did the manager reach out to apologize.
Usually we only complain about the short sighting of the CEOs that prioritize short term stock gains over long term prosperity, but that also is just a specialized case of the success vector misalignment
If OP is meeting deliverables and did this on the side, this is a great example of a story that would make me inclined to hire.
I was trying to say that the employee tried to do what’s good for the company. However that backfired as his personal, and his manager’s, success vector are not aligned with the company’s. Hence although he did something that benefited the company it hurt his and his manager’s performance (exactly because they are not aligned with that of the company’s).
Hope it clears things up.
Helpful in the right light, but equally not helpful in another.
Every team needs slack. Bad managers hate slack. Good managers hide slack.
The only problem is that the database code might need maintenance, specs change or situations change, and the cost is _not_ 25m, but way more per year than anticipated. So much so that after a few years, a new batch of engineers start eyeing writing a new database system to purportedly save money...
This is a huge ROI. Borrowing $25m costs about $1.25m/yr so you're winning even with no upfront costs
I suspect there's now a team of technologists maintaining, securing, and operating LedgerStore.
Hosted services doesn't reduce maintenance costs to zero
With LedgerDB, they now have to do both.
Opportunity cost is a cost, too
Imagine trading 5 headcount full-time to manage the 1T+ fully custom database on an ongoing basis when they could have just used DynamoDB and have been done with it.
Or better, having to engineer a new feature that already existed in DynamoDB and just losing money at that point.