Teradata is a lot faster for interactive workloads than Databricks.
PS: I agree there's no comparing on Databricks vs Snowflake/BigQuery.
Teradata is a lot faster for interactive workloads than Databricks.
PS: I agree there's no comparing on Databricks vs Snowflake/BigQuery.
Good luck on your budget trying to scale up shared-nothing database and making it scale up and down based on workload without downtime.
You can achieve significant speedup resembling shared-nothing databases by pushing the data close to the query using caching. Snowflake does it out of the box as it maintains table metadata. Databricks can do it too, but you have to be careful and it sucks.
Teradata was faster with no user-facing tuning vs tuned Databricks. And if you can pay for Teradata you may as well use it.
Self hosted HDFS will work more like shared nothing.
In shared-nothing the data "lives" on the compute nodes, so you can't willy nilly add or remove nodes. The data would either get lost if you removed more nodes than what is necessary for triple replicated redundancy. Or if you added nodes you'd have to wait before the data gets rebalanced to those new nodes, resulting in massive reshuffling of everything.
Keep in mind that shared-nothing clusters are most likely long lived and have multiple large datasets sitting on them, so by adding nodes you will start to shuffle everything around.
With shared-disk your compute cluster is only a single use for one dataset and you don't lose data as you scale cluster up and down.
There is a hybrid architecture which will give you both, but doesn't exist out there: Shared-disk with active caching. Ie. giving you control over which tables will be pinned down on your temporary compute-cluster. That will give you performance of shared-nothing but with convenience of shared-disk temporary compute clusters.
And having control over the cache sucks as hell. You can't pin down a table to reside in compute node disk cache when you know it will be used often.