Like this: https://use.expensify.com/blog/scaling-sqlite-to-4m-qps-on-a...
Like this: https://use.expensify.com/blog/scaling-sqlite-to-4m-qps-on-a...
Sqlite's own advice:
> If your data will grow to a size that you are uncomfortable or unable to fit into a single disk file, then you should select a solution other than SQLite. SQLite supports databases up to 281 terabytes in size, assuming you can find a disk drive and filesystem that will support 281-terabyte files.
> Even so, when the size of the content looks like it might creep into the terabyte range, it would be good to consider a centralized client/server database [over SQLite].
I don’t think you’d be able to fit this much storage into a single machine, especially not for a few thousand a month and SQLite wouldn’t be appropriate for this use-case.
> 1:1 replication Depending on the amount of writes could be a ton of extra disk and a bucket for network cost
Except network cost there's no extra disk required. It's just broadcasted writes consumed on the other hand.
These boxes are not dumb JBODS. They support their own replication/backup subsystems, so everything is transparent.
Keep scaling and eventually vertical integration ends up looking like a Soviet style planned economy. Your remote mining town needs some way for people to get soap etc so you open a store with it’s own supply chain etc etc.
But even at the 30% marketshare the iPhone has been dealing with these issues for a while. They just can’t buy 200 million volume buttons or whatever off the shelf. Now imagine what happens if one of their suppliers would fail days before the phone launches. They have such tight integration not just with the manufacturer’s process but also their finances because Apple simply can’t get replacements at scale on short notice. And remember that’s at ~30% market share, it just gets worse after that.
Tesla does some stuff in house because they can, but they just didn’t have any options when it came to scaling battery production. They hit the limits of what the market could supply without getting directly involved.
Suppliers don’t want to purchase a bunch of equipment and scale their manufacturing capacity for a contract that could go away at any time. So when Apple shows up they are going to price that risk in to their bid or sometimes just outright say no. However, by Apple buying that equipment those risks suddenly get reduced.
Apple on the other hand has huge cash reserves and isn’t that price sensitive. Plus they want to avoid all the issues bring yet another component in house.
https://www.supermicro.com/en/products/system/1U/1029/SSG-10...
I've never done procurement so can't really speak to recommendations. I just see that the products apparently exist.
There's also 32 drive 1U JBOF enclosures for more expansion:
https://www.supermicro.com/en/products/system/1u/136/ssg-136...
1PB in a rack with spinning rust + flash buffer has been easy for years now.
Lots of orgs fail to turn money into talent and then talent into products.
It just takes one bad hire at senior level and suddenly your cloud is a vmware install where all machines are boot off network disk, and contention makes the entire thing fall over.
[0] https://www.sqlite.org/releaselog/3_33_0.html
[1] https://www.sqlite.org/limits.html (#12)
Eventually all programs will be able to read email.
Plus size is only one limit, you would be limited to 1 write every few milliseconds. My napkin maths estimate is that there are at least 1-2m writes per hour going into this thing, so probably 300-600 writes / second (Average) and maybe over 1k writes/second peak. We are going to fall over here!
Not sure why some people seem to have a viwe of "There is no scaling problem that can't be solved with a sufficient enough number of SQLite databases".
1 is a scalable, managed, highly available service, with economies of scale the other is a fixed size, capital expenditure with fixed performance, limited DR, requiring a couple of SRE/DevOps and colo
There is also the will it always work question
https://sqlite.org/lang_attach.html
'Transactions involving multiple attached databases are atomic, assuming that the main database is not ":memory:" and the journal_mode is not WAL. If the main database is ":memory:" or if the journal_mode is WAL, then transactions continue to be atomic within each individual database file. But if the host computer crashes in the middle of a COMMIT where two or more database files are updated, some of those files might get the changes where others might not.'
Most of the novel work in LedgerStore is probably around managing the headaches of distributed storage, not the persistence layer.
A reasonable number for one server is about 32-128 TB, and 1.7 petabytes with some redundancy fits nicely in ~30 servers with a decent distributed database.
How do you detect restored but bit flipped data ?
sqlite3 /path/to/db
sqlite> PRAGMA integrity_check;
See SQLite3 documentation: https://www.sqlite.org/pragma.html#pragma_integrity_check- Client/Server applications (Check)
- High-volumes (Check)
- Large datasets (Check)
- High concurrency, particularly for writes (Check)