But if you got 5 TB of data, that needs to be in a SSD drive, then please tell me how I can get that into 1 single physical database.
But if you got 5 TB of data, that needs to be in a SSD drive, then please tell me how I can get that into 1 single physical database.
Now you don't have to shard. More info on how we accomplish distributed transactions. https://fauna.com/blog/distributed-consistency-at-scale-span...
Partitioning the dataset isn't really novel these days (Cassandra, Riak, Mongo et al do the same of course), but what is a significant difference is that both Spanner and FaunaDB implement ACID transactions distributed across partitions. It no longer matters for application correctness what partition key you choose if you can involve any arbitrary set of records in an single transaction.
> Whenever I've heard the term "sharding" it's usually in reference to the application-level sharding described in the article.
I wanted to drop a quick clarification note here. In the article, I used the term "sharding" to refer to both application and database level sharding.
For anyone that's looking at sharding as an option for scaling, we're always happy to chat and help point you in the right direction. My email's ozgun @ citusdata.com
If you're looking databases that come with built-in sharding, I'd definitely check out Citus (then again, I'm biased): https://www.citusdata.com/
if all you need is a lot of data in a single database, there's basically nothing except for money between you and your goal. JBODs full of SSDs coming into a single machine via SAS will get you into petabytes, just with commodity hardware you can order from amazon.
i'm expect IBM could sell you a mainframe that'll do it for whatever capacity you care to name.
if you're a reasonable person, you buy 6 1TB samsung SSDs and stuff them into a single 2U case and you're done.
Let's not pretend there is anything reasonable in this setup.
http://www.dell.com/en-us/work/shop/povw/poweredge-r930
Some of us do need to shard for sure though (I have multi petabyte data sets).
Also, operations on a such huge data set can be really painful. Think how to backup a DB like that safely, or how to update the engine.
Some slides (little old, 2014) about a huge postgres instance serving as a backend for leboncoin.fr (main classified advertising website in France).
https://fr.slideshare.net/jlb666/pgday-fr-2014-presentation-...
Basically, they bought the best hardware money could buy at the time to scale vertically, they, in the end, run in some issues and started thinking about sharding this huge DB.
I have a workload that runs close to 1.2million TPS for hours at a time and needs less than 100 millisecond response times at the 99th percentile. That uses more than 1 box and sits (replicated) in RAM.
However, 5TB of data really _isn't_ that much on modern SSD's. You can fit a sizable chunk of that in RAM on a decent server, so you probably _don't_ need more than one box.
I have 5TB of data that needs to sit on an SSD is, to be honest, a really poor performance metric. If you are genuinely specing out hardware and a database a better statement would be:
"I have 5TB of relational data, with a pareto distribution for access, at a peak of 100K TPS". Then we can start talking about what solves the problem.
Perhaps the title is click bait, but at the time I was meeting with a lot of users looking for someone else's problems.
5TB could still easily be single server territory. It depends more on the queries.
My point is just that some workloads are better solved with (some) vertical scaling first.
or did you mean just use a large capacity RAID setup? that will probably work fine for a lot of situations but it can expensive and introduce more latency for certain types of operations (but that might not matter, depends on context).
https://petapixel.com/2015/08/15/samsung-16tb-ssd-is-the-wor...
Drive capacity in a server is not limited to the size of a single drive. You can build a raid array any size you like by simply adding more drives.