Migrating Uber's ledger data from DynamoDB to LedgerStore
uber.com
uber.com
Like this: https://use.expensify.com/blog/scaling-sqlite-to-4m-qps-on-a...
I don’t think you’d be able to fit this much storage into a single machine, especially not for a few thousand a month and SQLite wouldn’t be appropriate for this use-case.
> 1:1 replication Depending on the amount of writes could be a ton of extra disk and a bucket for network cost
Except network cost there's no extra disk required. It's just broadcasted writes consumed on the other hand.
These boxes are not dumb JBODS. They support their own replication/backup subsystems, so everything is transparent.
Keep scaling and eventually vertical integration ends up looking like a Soviet style planned economy. Your remote mining town needs some way for people to get soap etc so you open a store with it’s own supply chain etc etc.
But even at the 30% marketshare the iPhone has been dealing with these issues for a while. They just can’t buy 200 million volume buttons or whatever off the shelf. Now imagine what happens if one of their suppliers would fail days before the phone launches. They have such tight integration not just with the manufacturer’s process but also their finances because Apple simply can’t get replacements at scale on short notice. And remember that’s at ~30% market share, it just gets worse after that.
Tesla does some stuff in house because they can, but they just didn’t have any options when it came to scaling battery production. They hit the limits of what the market could supply without getting directly involved.
Suppliers don’t want to purchase a bunch of equipment and scale their manufacturing capacity for a contract that could go away at any time. So when Apple shows up they are going to price that risk in to their bid or sometimes just outright say no. However, by Apple buying that equipment those risks suddenly get reduced.
Apple on the other hand has huge cash reserves and isn’t that price sensitive. Plus they want to avoid all the issues bring yet another component in house.
https://www.supermicro.com/en/products/system/1U/1029/SSG-10...
I've never done procurement so can't really speak to recommendations. I just see that the products apparently exist.
There's also 32 drive 1U JBOF enclosures for more expansion:
https://www.supermicro.com/en/products/system/1u/136/ssg-136...
1PB in a rack with spinning rust + flash buffer has been easy for years now.
[0] https://www.sqlite.org/releaselog/3_33_0.html
[1] https://www.sqlite.org/limits.html (#12)
Eventually all programs will be able to read email.
Plus size is only one limit, you would be limited to 1 write every few milliseconds. My napkin maths estimate is that there are at least 1-2m writes per hour going into this thing, so probably 300-600 writes / second (Average) and maybe over 1k writes/second peak. We are going to fall over here!
Not sure why some people seem to have a viwe of "There is no scaling problem that can't be solved with a sufficient enough number of SQLite databases".
1 is a scalable, managed, highly available service, with economies of scale the other is a fixed size, capital expenditure with fixed performance, limited DR, requiring a couple of SRE/DevOps and colo
There is also the will it always work question
https://sqlite.org/lang_attach.html
'Transactions involving multiple attached databases are atomic, assuming that the main database is not ":memory:" and the journal_mode is not WAL. If the main database is ":memory:" or if the journal_mode is WAL, then transactions continue to be atomic within each individual database file. But if the host computer crashes in the middle of a COMMIT where two or more database files are updated, some of those files might get the changes where others might not.'
Most of the novel work in LedgerStore is probably around managing the headaches of distributed storage, not the persistence layer.
How do you detect restored but bit flipped data ?
sqlite3 /path/to/db
sqlite> PRAGMA integrity_check;
See SQLite3 documentation: https://www.sqlite.org/pragma.html#pragma_integrity_check- Client/Server applications (Check)
- High-volumes (Check)
- Large datasets (Check)
- High concurrency, particularly for writes (Check)
Lots of orgs fail to turn money into talent and then talent into products.
It just takes one bad hire at senior level and suddenly your cloud is a vmware install where all machines are boot off network disk, and contention makes the entire thing fall over.
A reasonable number for one server is about 32-128 TB, and 1.7 petabytes with some redundancy fits nicely in ~30 servers with a decent distributed database.
Sqlite's own advice:
> If your data will grow to a size that you are uncomfortable or unable to fit into a single disk file, then you should select a solution other than SQLite. SQLite supports databases up to 281 terabytes in size, assuming you can find a disk drive and filesystem that will support 281-terabyte files.
> Even so, when the size of the content looks like it might creep into the terabyte range, it would be good to consider a centralized client/server database [over SQLite].
https://www.uber.com/en-US/blog/dynamodb-to-docstore-migrati...
Granted it's not quite in the same calibre as OpenAI/Claude, and the real test is when it is and they still release it.
It seems they need strong consistency for certain CUJs and then a lot of data warehousing for historical transactions.
It’s strange to me that they didn’t first convert their 2 table DynamoDB architecture into DynamoDB and Redshift architecture or similar. This is a pretty common pattern.
You can also look into “realtime data warehousing”.
Meanwhile, it’s so common, AWS built it into their product.
https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-dy...
It is breathe of fresh air to see that orgs are still building tools like that.
It turns out however that even when you use it as a distributed hashtable you still pay a huge premium.
https://www.uber.com/en-US/blog/bootstrapping-ubers-infrastr...
How so? It’s a pretty ubiquitous problem…
https://www.uber.com/en-AU/blog/migrating-from-dynamodb-to-l...
I'm guessing they know a lot about their costs, and you know very little. There's little value in insulting the team members like this.
It's an insult if you dismissively explain basic things to the folks working on the project.
I'm curious what makes you believe the OP doesn't know about cost? They might be director-level at a large tech company with 20+ years experience for all you know...
> There's little value in insulting the team members like this.
I'd argue it's not insulting to question a claim (i.e. 'we saved $6MM') that is offered with little explanation.
There is a lot of room in between, where cost estimates are more realistic.
Not in the US though. According to levels.fyi, an SDE2 makes ~275k/year at Uber. Hire 25 of those and you're already at $6.875MM. In reality you're going to have a mix of SDE1, SDE2, SDE3, and staff so total salaries will be higher.
Then you gotta add taxes, office space, dental, medical, etc. You may as well double that number.
And that's just the cost of labor, you haven't spun up a single machine and or sent a single byte across a wire.
And you can fudge the employee salary a mil or two either way, but the point is that spending that much on a team to build something isn't infeasible or even unreasonable.
Economies of scale help a bit with this for larger companies, so it's probably not quite double for Uber, but yeah, not too far off as a general rule of thumb. Probably a 75% increase on the employee facing total comp to get fairly close to the company's actual cost for the employee.
This is using existing features of Docstore which is Uber's own DynamoDB (sharded MySQL) which they seem to be using for almost everything.
They are structured and run like a tech company but imo they don’t produce a tech product.
This kind of comment could kill their stock :P
Uber said they were going to produce self-driving cars, so technically they are a failed tech company, which is now just a company. I am surprised they didn't get crushed by a competitor that didn't waste their capital on worthless stuff. So much for capitalism working. Uber should have gone bankrupt.
There is just no way that their database is actually technologically novel. Uber just doesn't have the expertise.
We don't know how much time and head count Uber committed to this project, but I would be impressed if they were able to pull this off with fewer than 6-8 people. We can use that to get a very rough lower-bound cost estimate.
For example, AWS internally uses a rule of thumb where each developer should generate about $1MM ARR (annual recurring revenue). So, if you have 20 head count, your service should bring in about $20MM annually. If Uber pulled this off with a team of ~6 engineers, by AWS logic, they should about break even.
Another rule of thumb I sometimes see applied is 2x developer salary. So for example, let's assume a 7-person team of 2xSDE1, 3xSDE2, 1xSDE3, and 1xSTAFF, then according to levels.fyi that would be a total annual salary of $2.3MM. Double that, and you get $4.6MM/year to justify that team annual cost footprint, which is still less than $6MM.
Of course, this is assuming a small increase in headcount to operate this new, custom data store, and does not factor in a potentially significant development and migration cost.
So unless my math is completely off, it sounds to me like the cost of development, migration, and ownership is not that far off from the cost of the status quo (i.e. DynamoDb).
Maybe you can do it a bit cheaper, e.g. with 4-6 people, but my point is that there's an on-going cost of ownership that any custom-built solution tends to incur.
Amortizing that cost over many customers is essentially the entire business model of AWS :)
In my experience you probably need a small team (6-8 people) to maintain something like this. Maybe you can consolidate some things (e.g. if your system has low on-call pressure, you may be able to merge rotations with other teams, etc.) but it doesn't go down to zero.
Also it’s not very clear from the original articles, what is the new total “cost of ownership” of this new refactored service. Like now they need to manage their own databases and the storage backing them. Or did i miss it?
My take on this is that most companies don't have the expertise to build systems like databases, and even if the costs would otherwise suggest such a development as desirable would be simply afraid of doing it.
They hired some out of school data scientist to do reports and they were doing crazy ineffective things with the tiny dataset. Wanted me to fix it for pennies tomorrow and I declined.
> I was doing some contract work for a small place that had a GCP Bigtable that was costing $11k+ per month for some reports that were based on data from a 375MB !!! mysql db into big-table for the reports to run.
Is a good example. It's just a badly architected system, and you'd have exactly the same problem if you were running the same thing on a massively over provisioned on premise db.
When I design computing systems, I am like "what's the point?", if there is nobody with a brain around to maintain it. The cloud allows companies to have services they could never ever maintain themselves long after I am dead.
We are going full Idiocracy from what I can tell.
You sound way too optimistic, but I am guessing you are still young.
For example, if a customer registers with the name of a deleted customer, which will resurface some "unfinished" transactions or rules associated with the older version of the "same" customer that haven't been properly deleted but appeared to be deleted for a while.
Also, in general, deletion is very difficult because money doesn't just disappear. You'd need some sort of compaction (think: Git squash) rather than deletion to be able to balance the system's books... but then you'd be filling the system with fake transactions...
From my experience from working with these kinds of systems, the typical solution is to label entities with active/inactive labels to substitute deletion. But entities never go away.
If you are not required to keep the information for more than X years, and you still keep it, then you have to provide it when it's requested.
If you didn't keep it, then it can't be used against you.
If you delete it after it was requested, then you are in trouble.
I tried same for ready to eat meal everyday to save me from potential kitchen disasters but sadly numbers didn't work out.
Turns out it costs more to DIY when people start quitting because they have to work during a holiday.
For our risk profile this was more than enough time to migrate off any AWS' proprietary technology.
That makes it worth less to avoid exposure.
You’re always dependent on your infrastructure. Even if you have nothing but everything hosted on a bunch of VMs, it can take years and millions of dollars to migrate.
No, just use Terraform and Kubernetes is not the answer.
The typical enterprise is dependent on depending on the source between 80 - 120 SaaS products - ie outside vendors.
I'd assume it takes fewer millions to migrate your own tech stack from AWS to somewhere else than it takes to migrate from AWS proprietary solutions. Is that reasonable?
And you’re dealing with your PMO department, project managers, finance, security, contract negotiations, retraining your ops department…
And you know that Aurora MySQL instance that was suppose to prevent “lock in”? I bet you someone somewhere in your org thought about creating an ETL job and then said forget it and used “select into S3” to move data from MySQL into S3.
As a project manager trying to ship code so you can show “impact” to put on your promo doc, are you going to choose for your team to spend weeks to write an ETL job to prevent “lock in” or are you going to tell the developer to write that one line of SQL?
There are all sorts of choices you can make that will save time and money and ship features that actually deliver value instead of worrying about the boogie man of “lock in”.
And I really hope that there was some better technical reason than just saving $6 million dollars a year for a multibillion dollar company to go through the migration.
But the main question is, once you do all of this work and spend time to be “cloud agnostic”, does it add business value?
In the case of Dropbox, it made sense to move from the cloud. In the case of Netflix, they decided to move to the cloud.
But you can’t stay completely “cloud agnostic”.
Let’s take a simple case of using Kubernetes and building the underlying infrastructure using Terraform.
The entire idea behind Kubernetes is to abstract your infrastructure - storage, load balancers, etc.
But eventually, you still have to deal with what’s underneath. I used AWS’s own Docker orchestration service for years - ECS. But I just learned Kubernetes last month.
I still had to know how to troubleshoot problems with IAM permissions, load balancers, view CloudTrail logs for permission issues, know how the underlying storage providers worked, make sure I had the right plug installed for K8s to work with AWS’s infrastructure etc.
Once I got all of that figured out, then I could go through the tutorials and mind map the difference between ECS and AWS’s Kubernetes implementation - EKS.
But I had years of experience with AWS. I could have never easily troubleshoot the same types of issues with Azure’s or GCP’s version of K8s. Now multiply that by an entire department.
Once everything is configured correctly, a developers experience would be the same across environments
Migrations at scale are always a pain from one system to another.
Source: I worked at AWS in the Professional Services department for three years. I’m mostly a developer and I dealt with the “modernization” side of “lift and shift and then modernize”.
It could be worse. It could be better.
It costs me $10M to run something every year, it now costs me $4M to run something every year. I have $6M in my pocket every year now in perpetuity. Compound that annually with the assumption that I maintain or increase top line revenues and thats pure extra profit.
Note - I admit, all of this ignores two key things (a) we dont know the engineers salaries who built this and (b) we dont know the ongoing maintenance costs.
That is because an investment is a one-off, so it's actually worth $1, but the savings are recurring, so they are worth the same number of years that a company's profits are valued. Depending on sector and investors' beliefs in the future of companies, this factor is typically in the 5-20x range. That means that $1 of savings is well worth at least $5 of investments.
Factor in anything you want!
Sigh.
So yes $1 of savings is worth more than $1 of spending.
if you assume they wouldn't have had anything else meaningful to work on during that time to save money, then you have a different problem in the company. $6M seems like the value 1 engineer can drive in a company at the scale of Uber
But if we’re being honest, there isn’t actually any meaningful quantification of engineering time to understand return on investments at this level (not to say there’s none, but it sure does get wish washy). Corporate and engineering strategy isn’t so carefully weighed, and to believe otherwise is to fall victim to the pseudoscience that is software estimation. You just have to estimate directionally if a given proposal has you heading in a better direction in the long term, pursue that, and course correct along the way.
Put another way, the end state justifies the means and resourcing. It’s rarely possible to fully understand either the costs or benefits with much accuracy up front. You slowly put more resources into projects that show promise, and revoke them if the projects do not appear to be heading in a value add direction.
You don't need to, but you 100% should. "Opportunity cost" (cost of not doing something) is real.
This is the problem with all refactoring/migration projects. It's very easy to get a lot of people to agree a company should migrate from Node to Go or Monolith to Microservices (or to clean up a mountain of tech debt), but it's much harder to justify the time it takes away from building things your users care about.
One can build a great career working only on key, promising initiatives that never amount to any value in the end. By the time it's clear the project lost money outright, you are on to something else.
People really struggle with large numbers in business I've noticed.
E.g., you can’t lay off an entire SRE team and have nobody on the on-call rotation. If some of their project work is cost control that is basically free cost savings.
No way they have 1 trillion transactions right?
Some records for the customer, some for the driver, some for the restaurant...
Eg you might have a record for each stage of the meal. When it's ordered, when it's cooked, when it's delivered, etc.
Does it have to be 1-dimensional? Depends exactly what payments is. There are refunds, discounts, paying e.g. drivers. There are also things like monthly subscriptions people can subscribe to for discounts / unlimited uses. Lots of things add up.
This seems low, off the bat. 15 years of Uber, 9 years of Uber Eats.
But even just looking at my most recent trip with Uber, there are 7 different records visible on the receipt. Not including backend recordkeeping that isn't exposed to the user (driver payments, driver loan repayments, revenue recognition, internal fees/records, etc).
Total trip amount, Trip fare, Booking fee, Tip, State fee, Payment #1 (trip itself), and Payment #2 (driver tip)
Now consider Uber Eats where there is (at least) one record for each item in an order...plus tax, tip, etc as always.
Then consider things like wait time charges, subscriptions, split charges, pending charges, chargebacks, refunds, disputes, blah blah blah.
An average of 10 records per customer transaction seems entirely reasonable.
Its only team who propose alternative they have to justify rigorously how come they differ in conclusion.
Yeah until those bill come, They would consider alternative
AWS has datacentres around the world, including multiple locations in the EU.
Schrems II prohibits transfer of personal information to companies reachable by the CLOUD act.
I think it will be very complex task to run MySQL for 1PB 1T transactions..
"The metaverse division has now lost more than $45 billion since the end of 2020"
Your compensation for your work is your salery. So I would say that it's fair that the actual risk taker is benefiting from the potential rewards?
But I guess we first have to agree on "who" we are taking about - is it the company itself or the owner / shareholders ?
Back to your question, yes that could happen in several different cases. But of course the risk/benefit is not split 50/50 (nor 0 risk, 100 upside, as you said), in reality the future outcome depends on both internal and external events.
Even the richest(?) man in the world was relatively close to loosing it all;
Musk, who had $200 million in cash at one point, invested “his last cent in his businesses” and said in a 2010 divorce proceeding, “About four months ago, I ran out of cash.” Musk told the New York Times https://www.cnbc.com/2017/04/27/the-crucial-decision-teslas-... https://archive.nytimes.com/dealbook.nytimes.com/2010/06/22/...
You wouldn't want all your earnings to be in stocks, you want liquidity. For example investing your earned money into a public company, or buying food.
what are the 'trillions' here?
^ also that translates to ~1000 transactions per second with some assumptions; have never understood why they care so much about infra scaling
1000 tps is like 1 box
> LSG promised shorter indexing lag (i.e., time between when a record is written and its secondary index is created). Additionally, it would give us faster network latency because it was running on-premises within Uber’s data centers.
https://www.uber.com/en-AU/blog/migrating-from-dynamodb-to-l...
1. https://www.uber.com/blog/how-ledgerstore-supports-trillions...
2. https://www.uber.com/blog/migrating-from-dynamodb-to-ledgers...
Saving $6M is key information that makes this story interesting. It’s buried all the way at the bottom of the first blog and is completely missing from the second blog which focuses specifically on the migration
However that appears to be defunct now
Because they were never submitted? I looked for the first one, it doesn't seem to be on HN.
"Uber Migrates" (beginning: company that I'm interested in does something) "1T records" (middle: that's a lot of records; I wonder what happened) "from DynamoDB to LedgerStore" (hmm, how do they compare?) "to Save $6M Annually" (end: that's a good chunk of change for me, but was it worth it to Uber? Why did it save that amount? Let me read more)
It's a simple and engaging "there and back again" story that paves the way for a sequel.
Versus:
"How LedgerStore Supports Trillions of Indexes at Uber" (ah, okay, a technology supports trillions of indexes. Moving on to the next article in my feed)
"Migrating a Trillion Entries of Uber’s Ledger Data from DynamoDB to LedgerStore" (ah, a big migration. I'm not sure who did it or whether anything interesting came of it, or even whether it happened or is just theoretical because of the gerund, and moving one trillion of something is cool but not something I probably need to read about right now, so let's move on)
YMMV. Some probably prefer the more abstract/less narrative titles, but the first one is more of an attention grabber for me.
https://duckduckgo.com/?q=blazing+fast+rust
Now I'm not an expert in either rust or go. But I know my deductive meme logic:
1. Uber's solution is not blazing fast
2. They are a Go house
Then the meme implies:
3. Their solution is slow because they did not use rust!
Q.E.M. (Quod Erat Memonstrandum)
My point is that a random dev running a pretty plain adblock (aren't we all?) simply cannot view their post. This is down to uber, their practices, an external developer and how uber create their blog (they don't just have the content in the page). If I'm not a special case with extremely weird luck, a bunch of devs seeing links to their posts will open them and not see any actual content. They will then, I assume, be less likely to upvote them.
Given that they are seeing problems with posts being upvoted this seems somewhat relevant.
You are running software that is blocking content you want to read. That is my point.
If I put on blinders and then complain I can't see your stuff, that's my fault not yours - regardless if your stuff is good or the worst annoying spam ever. If I want to see it for some reason, maybe I should take off the blinders
Yes. It's my point too. I am running very standard software for a dev and it is stopping their dev blog posts being visible.
> If I put on blinders and then complain
I'm not complaining. I'm explaining, given the evidence I have, why they may be seeing poor results on HN. If I'm not alone (and since I have no custom setup designed to keep our their blog posts that would be a surprise) then there are other developers who cannot see their posts.
Like turning off JS and saying webapps don't work anymore.
If the bug is on the devs then the devs are to blame, for maybe expecting teh ads are loaded, or the tracking third party code.
The project I am working on works with ad blocker on. Also we had issues with users that had a spellchecking extension active, it would create a ton of hidden markup on a contenteditable element, and we made code to handle the issue instead of having many tickets to our support complaining and we telling them that is their fault for using a popular extension.
edit - it doesn't have to really be blocking the actual post here even, if their loading code breaks when some other tracking code doesn't run, that could explain it.
Submitters: "Please submit the original source. If a post reports on something found on another site, submit the latter." - https://news.ycombinator.com/newsguidelines.html
Usually to aim for a significant promotion.
“Designed and built homegrown system to save $Xm! Give me promo, bro?”
Just so happened to ignore that it took X+Y additional to build. Also it will probably be going to the G graveyard in a few years.
If you're truly paying engineers, project managers, etc $500k a head, it dramatically undermines the financial cost savings.
It very well might be the case that "We spent $25m of engineering resources to save $6m annually".
[0] https://www.levels.fyi/companies/uber/salaries/software-engi...
[1] https://www.levels.fyi/companies/uber/salaries/software-engi...
Imagine trading 5 headcount full-time to manage the 1T+ fully custom database on an ongoing basis when they could have just used DynamoDB and have been done with it.
Or better, having to engineer a new feature that already existed in DynamoDB and just losing money at that point.
I made a config change to our AWS instances and projected approximately $10MM/year in AWS costs savings (pre-savings).
My boss asked me "Who told you to do this? We need to focus on $project instead". I found another team and transferred out. 3 months later there was a big fire drill about AWS costs and they took my 1-pager and executed it. Didn't get any credit in the shipped email nor did the manager reach out to apologize.
Usually we only complain about the short sighting of the CEOs that prioritize short term stock gains over long term prosperity, but that also is just a specialized case of the success vector misalignment
If OP is meeting deliverables and did this on the side, this is a great example of a story that would make me inclined to hire.
Helpful in the right light, but equally not helpful in another.
Every team needs slack. Bad managers hate slack. Good managers hide slack.
I was trying to say that the employee tried to do what’s good for the company. However that backfired as his personal, and his manager’s, success vector are not aligned with the company’s. Hence although he did something that benefited the company it hurt his and his manager’s performance (exactly because they are not aligned with that of the company’s).
Hope it clears things up.
This is a huge ROI. Borrowing $25m costs about $1.25m/yr so you're winning even with no upfront costs
I suspect there's now a team of technologists maintaining, securing, and operating LedgerStore.
Hosted services doesn't reduce maintenance costs to zero
With LedgerDB, they now have to do both.
Opportunity cost is a cost, too
The only problem is that the database code might need maintenance, specs change or situations change, and the cost is _not_ 25m, but way more per year than anticipated. So much so that after a few years, a new batch of engineers start eyeing writing a new database system to purportedly save money...
> Uber is famous for NIH syndrome
If I were doing this, I would be looking at data warehousing systems. 1.7PB of, say, Parquet files in S3 is not terribly expensive. 1.7PB of Parquet files in on-prem or collocated object storage, even replicated a zillion times, is quite cheap. And quite a few companies and open-source projects are currently competing aggressively to provide awesome tools for querying that data.
The hot data would fit on basically anything — the choice should be about robustness and barely even consider cost per TiB. Datomic got written up recently and seems credible for this type of application. FoundationDB is bulletproof. Postgres could probably handle it without breaking a sweat, although active/active replication isn’t free. Heck, writes straight to a warehouse with a cache in front to help with reads seems credible — Uber rides rarely go for longer than a couple hours, and back-of-the-envelope math suggests that the total data rate is maybe 50GB/hour. An entire day of data for an entire country would fit on a single very ordinary commodity server, and the live data for the entire world would fit on one mildly beefy server. The indexes involved sound straightforward.
Old data could probably live at lower cost in a data warehouse, but then developers would have multiple systems and namespaces to deal with in order to query on transactions.
What are you basing this on, banks have less real time constraints when it comes to certain data but thats rare. You will almost never see a bank, processor, etc. archive their system of record data.
This seems like an odd requirement to me. What are the use cases beyond a few minutes past the end of the ride?
- Reconciliation of Uber's finances. This definitely needs to be correct and consistent and maybe even auditable, but it's also done at the end of any given billing period. A daily roll-up would be sufficient.
- Showing each driver and each rider their own history. This ought to be fully correct (especially for drivers), but it's a massive double-edged sword. Uber's app, for example, appears to be entirely missing the "please forget my ride history" button. And this isn't surprising: Uber apparently stores it in a ledger, fully indexed, complete since at least 2017, that is cryptographically immutable. Is this actually a good thing? What happens when a privacy regulator tells Uber to expunge, anonymize, or pseudonymize old data?
- Handling chargebacks? Having a high quality record going back 180-365 days seems useful. But the cost to Uber of occasionally losing a record is very low and is barely more than proportional to the probability of loss of a given record.
So I don't get it. If I were running an operation like this, I would prefer not to have a fully immutable ledger full of personally identifiable location, usage, and transactional data.
Exactly what he the OP said, authorization. Card networks cap latency of authorization requests (the time the merchant submits a request, to when it makes its make through the merchant processor, networks, issuer processor and back.)
Yes, which they already had as DynamoDB only held 8 weeks of data. Presumably all payments to drivers were calculated monthly and turned into invoices. Any corrections would end up as adjustments on future invoices. Seems pretty normal.
Promotion-driven development. I suppose better than blog post driven development, but marginally so.
I find DB development quite interesting, and I find Uber's core product quite not interesting (from an engineering perspective).
So as an outsider without any financial stake in the company, please keep writing about databases!
Someone once called Airbus planes "Sun servers with wings". All businesses today have an IT department and those build around the idea of real-time scheduling done by computers are compute/storage businesses first. That's why Uber was able to launch Uber Eats or e-bikes. I am not surprised they have a bunch of busy devs building cool stuff. Not all of it is ego-driven. I liked their logger for Go last time I used it.
The architecture where all the magic lives in the database layer and the services treat the DB like this magical thing that synchronizes across multiple DCs for you, and takes care of every complicated part of engineering a distributed system is a "now you have two problems" situation. It's a massive complicated system that will fail in novel ways, and you don't have a large community of other users that you can fall back on for expertise.
That is one of the most important points of consideration when I am being asked to asses new tech. If the project is not actively supported by the community, I don't recommend it. I just can't expose the client to risk just because there's a new shiny thing.
(I understand there won’t be any significant business without enterprise sales. But that’s not what I’m looking for at this stage.)
A good chunk of B2B infrastructure products like this are developed using a “golden partner” model. The first customer (or few) gets a free or reduced cost license, the developer gets a real-world scenario with real data to use to figure out what the minimal functionality actually is to be a marketable product and to work out bugs. This arrangement frequently requires a preexisting relationship and trust between both parties.
A “golden partner” model makes a lot of sense, thanks.
No, you don't. There are many established storage solutions out there. If you're in the market for one, you can easily fill days, weeks or months vetting those. So, why would you bother dealing with a sales rep from a random one you never heard of before, and isn't used by anyone. You don't even provide any details on what makes it different or better from anything else out there.
I guess what I’m trying to say is that I was hoping that someone with a write intensive workload would want to spend some time evaluating a product built specifically for that. But perhaps I’m wrong? Even if your workload was 99% writes you’d rather go to some established player (e.g. MongoDB) with a product optimized for 50/50 read/write?
Again, it's not clear to me exactly what it is you're doing that's any different from the plethora of existing off-the-shelf solutions.
You're saying that you started this project/company because you were looking for a solution to a specific use case (write-intensive workloads) and existing options didn't work - can you expand on that? Can you create a chart, for example, that lists out the specific things that Haystackdb does and alternatives don't? Presumably, if you optimize for write-intensive workloads, there are some drawbacks when it comes to reads - no? Or maybe storage? That's good to highlight.
What you need are whitepapers/blog posts/youtube videos/talks at conferences/etc. that highlight the technical details of your solution, because you're trying to get technical people interested in your product to the point where they will invest time to learn more.
From pricing: “$0.2 per million writes, $20 per million reads”. The typical cost profile is $2 per million read/writes, or even more for writes.
> HaystackDB is designed from the ground up for write intensive workloads
Okay.
> so it’s much more economical than existing off-shelf-solutions for that type of workload.
That's a leap in logic. Just because you designed it with this workload in mind, well, doesn't automatically mean that it's any good for this workload (or any workload). If solving a problem was as easy as declaring "I will design my solution from the ground up for this problem", then we'd all live in peace and harmony. So that's what people are asking you here: how do you make your DB "much more economical" for that type of workload? What technology, what ideas have you had to make it possible? If you don't want to reveal that, then you need proof that it's better than the competition, not a declaration, that it's better than the competition.
> Is that not clear from the landing page?
It's clear that you want to market your solution as something good for write-heavy workloads. Why should we believe you've done a good job designing your solution?
> From pricing: “$0.2 per million writes, $20 per million reads”. The typical cost profile is $2 per million read/writes, or even more for writes.
Who knows how you came up with pricing? Perhaps you're betting on your customers being stupid and not realizing that taking a 10x hit on the price of reads will lose them (and earn you) more money in the long run. After all, what good is writing to a DB if you never read from it...? Or perhaps it's some kind of promotional / loss leader pricing that will change soon in the future. In any case, it's, again, not proof that your solution is adapted to the customer's problem.
No worries. I appreciate you taking the time.
> you need proof that it's better than the competition, not a declaration, that it's better than the competition
Fair point. I realize I’ll need that before making any sales. But I was hoping to get a few leads from the contact form without it.
> Perhaps you're betting on your customers being stupid and not realizing that taking a 10x hit on the price of reads will lose them (and earn you) more money in the long run. After all, what good is writing to a DB if you never read from it...?
No it’s not a malicious trick. There are use-cases where most records will never be read back. For example, if you go into the Uber app you can find a history of all your trips and you can click one and bring up a receipt for it. Most users will rarely if ever do that. So you end up writing many more receipts to your database than what you’ll ever retrieve.
The marketing byline you have on your landing page is clear enough, but nobody will take that seriously without a deeper technical description.
When I read it, I assumed you wrote some code to move data in and out of lower-cost S3 or Glacier storage tiers because you don't control storage pricing and you run on top of existing public cloud infrastructure. Maybe I'm right, maybe I'm wrong - but if I'm looking for a solution, I need to assess whether I should invest time and effort to do a deeper dive, and that's the box I would put you in, without any more detail.
Anyway, good luck. Hope it works out.
try connect to the respective people at said teams via LinkedIn and ask feedback
1. It is a bit unclear to me when I would use Haystack. The main advantage seems to be cost cutting. It would be nice to see some realized examples of this.
2. When competing for price, you may look like the cheap, and thereby untrusted alternative. There is a risky business paradox here, for which I am sure a fellow HN poster will supply the name: you charge less, therefore you make less, and you will not be able to sustain the service, making me not want to spend money.
3. Have you tried looking for companies that may actually need this solution? Have you tried contacting them directly?
2. True. One reason I haven’t priced it ridiculously cheap is to avoid this judgement, and fate. With this pricing I won’t necessarily have a smaller profit margin than competitors. The cost advantage comes from a smarter architecture. Any ideas on how I can communicate that would be greatly appreciated.
3. I used to work for one that needed it. I’ve also interviewed at one that had the same problem. A bit hesitant to reach out to potential customers though before I have a solid product I can deliver. But perhaps I shouldn’t be?
Companies generally have to be suffering pretty badly to take a risk on changing their tech stack to something unproven. And the risk for you at that point is that they choose to spend 10x on consultants to implement some existing system instead.
The CTO needs to trade off the opportunity cost of developing new features/existing maintenance against integrating an unproven product. How can you de-risk this for them? (Even just showing that you recognize that this is the case can help)
Maybe this is a time to "do things that don't scale". ie: offer to integrate it into their system for them (for at least some small part/pain point), and likely in parallel so that they can evaluate it without taking down the existing system.
Just my two cents.
I think you should invest some time into improving your landing page and maybe you may see some traction. A good resource for this which I've bookmarked is here(1). Hope that helps.
(1) https://www.indiehackers.com/post/my-step-by-step-guide-to-l...
To me, when I read the below, that just screams “save money”. But maybe I should do that conversation for the reader so to speak?
From benefits box: “Sometimes you need to index a huge amount of data, to accelerate just a few search queries. But building indexes and keeping them in hot storage can be expensive. HaystackDB builds only the indexes needed for sub-second query latency, across billions of keys, while keeping all your data in low-cost object storage like S3.”
HaystackDB: Swift Searches, Massive Savings - Index Billions, Store Smartly, Query in a Flash!”
EDIT: Or maybe it’s because reads are expensive? That’s a consequence of the write optimization. The idea is that potential customers will be doing 90%+ writes.
also is there a demo or some sort of technical whitepaper.
Structurally, you are a small entity trying to compete on cost with hyper scalar cloud providers and open source software. Most ISVs like you charge a ton of money for big problems very few enterprise customers have.
I think you need to find a specific use case where your product is a clear winner. Like 'HaystackDB is the best option for healthcare exchanges to use when receiving claims'.
The counter argument I guess is that developing your own data store in-house should be even more of a no-no, and companies do that. (One example is obviously Uber, but my previous employer is another example.)
Do you think the option to self-host the product would help tip the scale?
> What problem does your product solve that customers absolutely 100% need it?
To be blunt there is no such problem: you can always throw more money e.g. at DynamoDB. But if you have a very write intensive workload (such as the use-case described in the OP), then you can save 90% of that money.
Quite frankly, this is not gonna work. I manage a system with a very write-heavy workload (lots of small writes) and even though our writes far outpace our reads, this pricing makes your system about ten times more expensive than an RDS cluster.
There's no data about performance. There's no information on how or whether data is persisted to durable storage before a write is acknowledged. There's not even any information on how big keys or values can be. There's no public information on support.
When choosing a system like yours, my priorities are:
1. Data safety
2. Performance
3. Cost
... In that order. You've done nothing to educate me on 1 and 2 and your pricing isn't better than what you're seeking to displace.
When your product is a tool for developers, show up with hard facts about your product. Zero people (as you've seen) are even remotely interested in building a product on top of a system without knowing whether the system will hold up to their use case. And other than a very anemic FAQ section, you have no documentation at all, whatsoever.
> even though our writes far outpace our reads, this pricing makes your system about ten times more expensive than an RDS cluster
That indeed sounds off… Are you sure you’re comparing the total cost to that of an RDS cluster? I am aware that reads will be more expensive (due to the write optimization), but I was hoping most customers would make it back on cheap writes. Also the storage itself ($0.23 per GB-month) should be much cheaper than RDS.
At least I'm my case, the fundamental problem you're facing is that reads are just too expensive. Writes and reads tend to grow at the same pace in many products: there's a ratio that tends to stay the same as you scale. $20/million reads is just a _lot_. The ratio of writes to reads for your pricing needs to be 100:1 or more for it to make sense for me, but I'm more like 10-20:1.
> I guess I’m hesitant to put time into documentation and similar, if I can’t somehow find a steady stream of sales prospects.
This is part of why a database company is hard to build. You will simply not find anyone willing to give you money, because the alternative is going to be a solution your customers already know and understand and which is likely extremely mature. You're competing with Postgres and Mongo. You can't ship a database product that doesn't work: you're asking people to build on you for their storage primitive. If you fuck up, that's a business-ending event for your customer. You've either got to come to the table with an extremely compelling product ("I couldn't build my business without this") or you've got to show why someone should trust you over an established but somewhat more expensive alternative.
Correct. I bet Uber’s use case here is something like 1000:1. I’ve worked on systems that were over 1000000:1. That’s where HaystackDB makes sense.
> but I'm more like 10-20:1.
Then RDS is hard to beat.
Like a side-by-side example. Doing "work" on BigTable (show code examples) versus doing the same "work" on Haystack. Then show the specific metrics on how Haystack is cheaper/faster/better.
Seems you're in the vicinity of Lund, should be a 'science park' or similar close to the uni where you can find companies that have problems you could solve. Talk to 'incubators', 'accelerators' and the like there.
„how much talent is wasted on pointless things that help noone in the world while getting paid heaps for nothing”
We could accomplish everything if ppl stopped wasting time on pointless tasks.