LedgerStore Supports Trillions of Indexes at Uber
uber.com
uber.com
I'd like to see this broken out more. You replaced a managed service with a custom DB on top of MySQL. Surely there is a significant amount of engineering effort that was/is being spent to create/maintain this.
It's the dirty little secret of our industry.
When my last company transitioned to cloud, all of the on-prem folks had to additionally take on that responsibility. We then had to hire a bunch of new folks to ease the transition and manage the cloud pieces. And then there are new teams spun up for permissions and security. Your headcount requirements are not going to go down because of cloud.
As for those that like to call DIY systems "weirdware": homegrown systems are expensive to engineer, but cheap to run. We 10x'd our cost going to 3rd party visibility and feature flagging versus the thinly staffed [1] homegrown systems. And those 3rd party systems also required a ton of engineering to support, the migrations weren't clean, and they missed a lot of key features that we had to upstream to the third parties!
When SignalFx got bought out by Splunk, I've never seen an annual plan get re-wired so hastily. Top priority was moving to a new vendor, and that impacted every single team in the company.
The only real tangible benefit to cloud that I've seen is that a team can instantly provision hardware, database, etc. resources without much planning. That's it. And is that worth it? (I don't really think so.)
[1] Once built, half an engineering headcount per quarter to accommodate new features. Oncall rotation for one team that owned the systems. Our stuff was high resiliency and barely paged at all. No SEV1 or worse outages ever.
1) "homegrown systems are expensive to engineer, but cheap to run" misses the most important part - how much does it cost to maintain, and how much of a RISK is it?
Homegrown probably means you're depending on tribal knowledge from a few core people who if they leave you're hosed. That's a risk. You also have to train EVERYONE you hire to learn the homegrown systems, instead of hiring people with externally transferrable skills that can come in and hit the ground running.
2) all of the on-prem folks had to additionally take on that responsibility
Most transitions/migrations always end up with a period where you have both systems running. It's not surprising that at first everyone has additional responsibilities. It sounds like the "on-prem folks" either didn't have the skills to work in the cloud or refused to learn, and needed external hires. Well, that's one way to put themselves out of a job, and that would be the next step once the homegrown system is deprecated.
Obviously there's a reason why "not invented here" syndrome exists. There's a reason "build vs buy" is a complex discussion because of more than just buy costs. People always prefer the home-grown thing, DIY. Engineers especially. On this site, particularly. And, also, plenty of successful businesses exist by cutting out layers and going on a shallower stack ("your margin is my opportunity").
But at the end of the day, for the vast majority of businesses, companies, and software shops, managing their own in-house infrastructure is a worse decision than paying a cloud vendor.
there's so much complexity and tribal knowledge with cloud deployments that if tomorrow our cloud experts leave we're also very definitely hosed too, despite everything being documented thoroughly. I'm involved in a product that leverages a cloud-based metaverse system (Mozilla Hubs) that recently had organisational changes requiring us to change our hosting approach, and it's taking the better part of a year of work to understand its logic, for something that would have been a non-issue if homegrown & self-hosted.
Yep. Huge amounts of marketing funds are spent tricking decision-makers into thinking that off-the-shelf, plug-and-play solutions exist for their idiosyncratic business problems and schemas.
Sometimes such products exist, but they are so bloated with features to support the other 99% of customers. So you absolutely need at least one expert to unwind the complexity.
It's known that "greenfield" software development is much easier than "brownfield"/migrations. It should also be widely known that brownfield SaaS integrations are troublesome, because it involves enormous software+data work to link the existing in-house interfaces to the third-party interfaces. As a rule of thumb, there is no such thing as "off-the-shelf". You WILL have to hire and build and deeply understand your business processes.
And almost all business integrations are highly bespoke.
They can sometimes be very odd force multipliers in unexpected ways.
My favorite example is a home-rolled payroll applications specifically for sales people from over a decade ago. When I first arrived at the org I thought they were absolutely insane to have created such a thing. But its killer feature was that it allowed our org to pay out commissions rapidly (same day) and also allowed negotiating on a per-account residual commission basis as well as giving management the option to buy-out residuals. Off-the-shelf solutions at that time kinda-sorta did this but required either massive upfront $$$ and a ton of integration work. A plucky startup couldn't afford it. This was built, tested, and rolled-out in 45 days by 2 engineers and basically allowed them to poach top sales people because they knew they could get paid faster and had more flexibility on residuals.
Engineers make the same mistake all the time too: we love to say that "premature optimization is the root of all evil", ignoring that the full quote is "We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. _Yet we should not pass up our opportunities in that critical 3%_"
Yes! Maybe not for you or me, but it’s definitely worth it to tons of organizations where everyone including engineering was beholden to old school IT departments that take months to spin up a VM after pages of bureaucratic back and forth. The cloud moves it from an IT issue to department budgeting.
The other advantage of cloud was moving capex to opex which the finance people liked because *hand wave* something about amortization. Those are the two reasons most established companies moved to the cloud. It had little to do with the actual cost but how it was spent and who was in charge of doing it.
where everyone including engineering was beholden to old school IT departments that take months to spin up a VM after pages of bureaucratic back and forth. The cloud moves it from an IT issue to department budgeting.
My F500 company put a stop to that. There is new paperwork in place to requisition cloud resources. The bureaucracy will not be replaced.You end up wasting huge amounts of money on resources that are just sitting around idle (no your fancy automation to shut them down won't work because maybe that dev cluster is actually needed at 2am on Sunday. It's an org problem, not a technical problem).
But more importantly you give the cloud providers an excuse to destroy another square mile of pristine forest and build a giant 4 story grey box that wastes enormous amounts of already limited water and energy.
Virtual machines were the bureaucracy shortcut. People sounded exactly like you. In fact, you can go back about 30 or 40 years and each decade people said the same type of thing. E.g.:
"What's great about a mainframe logical partition is that we don't have to do a year of paperwork to get our own mainframe."
"What's great about Intel servers is that we don't have to beg IT for a mainframe LPAR."
"What's great about virtual machines is we don't have to... "
"What's great about the cloud is we don't..."
"What's great about SaaS is..." <- You are here.
PS: There's products starting appear in the market that "firewall" SaaS products to stop people like you. Not the hackers. You!
The bureaucracy will not be side-stepped.
Depends on how big you are. I'm one person and I do a lot of things with the cloud that I could not do if I had to rack servers (or even order dedicateds).
Same goes for small 5-10 person teams that are good with Terraform. I've seen some orgs punching waaaay above their weight given their size. Not possible without classic IaaS.
It's not a binary choice between cloud and physical racks in a data center. There is a wide range of options. Companies like Hetzner enables you to click a button and get a dedicated instance, while Digitalocean can provide you with a VM. Both options are cheap and doesn't require much time.
If you do this often, you probably also have an ansible setup to install the base software (like k8s or a simpler stack), you are ready to go within minutes. Just like you would have a terraform config for your cloud
It's like outsourcing (I don't mean offshore, I mean any kind). It can be a reasonable thing to do if you don't pretend you're getting the same thing for less money and really actually plan around that. But people often do pretend it's the same, and the results aren't immediately apparent. And when they're finally blatantly apparent, either the guilty party has cashed out and left, or people don't have enough perspective to point to the real root cause.
I'm not saying building your own infra is generally a good idea of course. It's just that beyond a certain size, for more businesses than you'd think, reliability has to be a core competency you invest in. And homegrown is a lot easier than it used to be because the offerings in terms of tooling and platforms are way better such that you can strategically craft your level of managedness for optimum cost/benefit on multiple criteria.
As a solo dev just being able to point my apps to firebase and not worry about it saves a ridiculous amount of time.
At Uber's scale they probably have to tweak Dynamo and other services so much they might as well bring em in house.
I'm a bit surprised we haven't seen a new AWS, something like Walmart Web Services.
Managed databases provide a lot of useful features, especially if shit hits the fan. I'm not saying you shouldn't use them. But the work you put into reasonable self-hosted services and reasonable managed services often scale at a surprisingly similar pace.
I still have to host it somewhere.
Plus you own the accounts, so you are not locked in.
Sync is the firebase kill feature, but most apps can live without it.
Firebase is ready to go without me thinking of the details.
If I'm using Flutter or React, I already have a nice client side sdk to use.
The 12th of never when I need to scale it I can switch to a different stack. Firebase functions allows me to do some server side data manipulation.
That said, if I was at work and needed to propose something I'd probably spend more time thinking of this. But for all my side projects firebase is my goto.
Let's just say we're having a competition to see who can crank out a basic crud app first. You're not beating me with Flutter and Firebase.
Even if you did, on premise self-hosting does not save from scaling costs or problems .
If you are successful with firebase , all you get is the bill, firebase will definitely scale without sweat or intervention, for most businesses (side or small) that is better tradeoff .
If you are successful with a self hosted solution , you are likely to be down when the peak traffic or hug of death hits that is when you need it most to work, the app will fail.
This is even assuming intimate expertise on how to scale quickly and having many done it many times before with no mis steps and no researching on what to do.
A professionally managed service at firebase level which is shared infra has already both the capacity and code tuned to scale for any one tenant without effort or lag .
Self hosted infra is optimized for cost and usage pattern of one small tenant, no matter how good your scale out skills and autoscaling code is , there is going be a lag between you and something like firebase which will never notice your spike and issues around it.
This is the crux of success behind multi-tenant IaaS and PaaS, the reason Amazon opened up their infra to the world originally.
All the way even to Amazon.com scale of SRE skills and budgets, hosting on AWS will always be better and cheaper for them than using isolated infra only for them given their load pattern fluctuates pretty heavily .
Amazon.com would still run AWS even if was losing money a bit ,because it would make them more resilient and their infra costs cheaper .
The best analogy I have heard is it similar to insurance, works better with more and more users with diverse usage and risk patterns
I find coding backends to be very boring. It's just not something I want to do during my spare time.
Managed infra is almost always easier to get going.
Generated CRUD interfaces with basics like authentication etc is pretty much easy these days, whether Supabase, Hasura, PostgREST style self hostable stacks or just Firebase, even collaborative data structure basics for CRDTs don't need to be built each time and are available out of the box.
Frontend can get repetitive too, a lot of components keep getting built and redone differently, interesting stuff which makes the tech be core part of the product differentiator such as say how WASM is used at Figma is the real interesting part in frontend.
Don’t know if you have heard but there are these brand new services called Google cloud and Microsoft azure. Not to mention all the legacy saas providers e.g. oracle that existed prior to aws.
It’ll bite the company in the ass eventually, or maybe it won’t, but it is what it is.
It's a tradeoff, but I don't think folks are keeping secrets around it.
I guess I finally understand what they meant by "no one ever got fired buying IBM".
I mean I’m here for the tech of course just curious from the cost benefit part.
Harkening back to Paul Graham’s “Beating the Averages” talk about Lisp - they’ve solved problems that their competitors are also trying to solve, which could give an edge.
From the valuation perspective, for that same reason, this is now an asset that can be acquired as a technology on its own - whereas using MySQL adds no intrinsic value.
You can then either write a new database that fits the use case, or take a hard look at the use case and change it to fit existing database systems.
Now as a programmer I LOVE the idea of writing a database or an operating system, it would be super leet.
When writing a new database system, you will make a lot of mistakes and a lot of bugs that you will no doubt encounter at inconvenient times after a lot of debugging and profiling.
This is built on top of MySQL which has been battle tested (still not where I would put any critical data) Technically that you're not writing the engine but with so much functionality in two layers above MySQL you are in essence building an engine for the engine.
The implicit assumption of "not invented here syndrome" is that it has been invented elsewhere already, which isn't always true.
this provdies room for doing real customization without having to build all the things from ground zero, and is likely to be more robust than trying to attach functionality fully on the outside like was done here.
in the database world it would be really nice to implement caching and sharding extensions inside the transactional envelope
The 6m in savings does not properly account for things like ramping up new hires on some custom database, maintenance (what happens in 5 years when whatever language you wrote it in needs to be upgraded to a new version, or some dependency), and a host of other things.
Yes the cloud is expensive, but the entire point of it is that you are offloading all of that underlying maintenance/feature work to a team that only does that all day every day and is very good at it.
This seems to be an in-house DynamoDB (sharded MySQL with Raft) which has been developed a long time ago, see: https://www.uber.com/en-US/blog/schemaless-sql-database/
But maybe I'm totally wrong here because they also use or used Google Spanner: https://www.uber.com/en-US/blog/building-ubers-fulfillment-p...
I think the issue is their engineering culture. They have always had far too many engineers. They are clearly over complicating a such a simple product.
Previously discussed on HN (294 comments):
As a commenter there put it - """The actual summary of the article is "The design of Postgres means that updating existing rows is inefficient compared to MySQL"."""
I think this is yet another approach to dealing with the cost of updating rows?
I'm curious about what it would cost you to run the same systems on Spanner.
(Sure, I hate Oracle too, but they actually did a pretty good job of both Java and mysql; i guess the 90s Borg for MS Gates and Satan for Oracle Ellison are less things now than they were)
What kind of cases. Since PostgreSQL wins in benchmarks I’m curious what’s faster.
Also, if you’ve designed your schema and queries around MySQL (InnoDB, specifically), you can get way faster queries. It’s a clustering index RDBMS, so range queries require far fewer reads thanks to linear read ahead, but only if your PK is k-sortable. UUIDv4 will absolutely tank the performance. It will on Postgres too, but under different circumstances, and for different reasons.
Benchmarking anything is fraught with peril. There are so many variables, and the benchmark itself can be quietly selected to favor one or the other. I’d encourage you to create identically-sized instances, with your actual schema, and realistic test data.
They are wrapping another database that solves these hard problems for them. This is a fine program to write. They will be able to hire other people to maintain it. Its not rocket science
Perhaps the lines are wrong?
As the article explains, it's not, at least from the merchant's side (from the customer's side it is). A CC hold tells the bank "hey guys, make sure that there will be at least X dollars available when we call in the hold", and the bank can either respond with "yes, hold confirmed" or "nope, declined" (say, due to a lack of funds, wrong CVV, whatever). When the hold is confirmed, and the customer makes another transaction at another vendor that would cause the amount of open holds + executed but not paid-off transactions to go above your limit, that transaction gets blocked.
And eventually, only once you as the merchant call in the hold, you will actually initiate the flow of money. Or you call in less than the held amount, and you'll only get that amount of money (usually seen in hotels and car rentals, where the hold can be significantly larger than the bill, to account for stuff like minibar expenses, room cleaning or damages to the vehicle). Or you release the hold, or it expires, and you get no money at all, and all you can do is try to create a new transaction and hope it gets approved.
Capture contains the final transaction amount and can either be higher or lower than the Authorized amount.
The odd thing is the Trip Service shouldn't know anything about the Issuer. That's the Payment Service's reason to exist (so Trip Service doesn't need to know about different payment methods). That the Trip Service knows how to talk to Visa (and Amex, and MasterCard, and Discover, and Google, and Apple) about Auth'ing a card, but doesn't know how to Capture the payment is the strange part.
My expectation is that the diagram is incorrect, and the TripService talks to the PaymentService to Auth the card.
The idea that Latin-derived words should have Latin-style plurals in English makes for pointless inconsistencies in the language.
We should use falcor! Netflix stops supporting it.
We should use functional programming! It became far too complex and we have abandoned it.
Okay, time for a new API. Let’s use vanilla express and do everything on hard mode. Some people never learn.
which you ought to be charging a pretty penny to do so.