I guess the scale of data here ~100GB is manageable with something like DuckDB but once data gets past a certain scale, wouldn't single machine performance have no way of matching a distributed spark cluster?
I guess the scale of data here ~100GB is manageable with something like DuckDB but once data gets past a certain scale, wouldn't single machine performance have no way of matching a distributed spark cluster?
They happen all the time if you work for banks, large finance companies or government.
It’s not just the 2TB databases - it’s the 100 analysts all doing it their own thing with that data at the same time.
This is my problem with databricks. It seems like in the course of selling their product they have taken the received wisdom of "do not run expensive and complex compute clusters unless absolutely necessary" and turned it into "it's fun and easy to run distributed compute clusters - everyone's doing it and you should too" regardless of how contextually appropriate it is.
Microsoft’s synapse is their «ripoff product» which tries to compete directly by offering the same thing (only worse) with MS branding and better azure integration.
I’ve yet to see Spark being used outside of these products but would be happy to hear of such use.
When it's more common to be manipulating lots of smaller datasets together in a way where you need to have more than 100GB of disk space. And in this situation you really need Spark.
If you can afford it you should be using the host NVME drives or FSX and relying on Spark to handle outages. The difference in performance will be orders of magnitude different.
And in this case you won't have the ability to store 64TB. The average max is 2TB.
We normally still would only need to process say, the last 10 days of user data to get decent recommendations, but occasionally it would make sense for processes running over the entire dataset.
Also this isn't that large when you consider binary artifacts (say, healthcare imaging) being stored in a database, which pretty sure that's what a lot of electronic healthcare record systems do.
A random company I bumped into has a 40TB OLTP database to this effect.