As a DS/DE, there's a lot to love (not all, but a lot). The easy provision of Spark clusters. The jobs API. DeltaLake (mostly). Easy notebooks (please don't create a prod system from these..). And Spark itself continues to improve, albeit in an increasingly crowded field.
But I've worked closely with BigCo SQL analysts on Azure Databricks, and their experience was terrible. For example:
- You cannot browse the data structure without an active cluster
- Starting a cluster can take ~5 minutes and, since you missed that moment, you may not submit your first query until 10-15 minutes.
- The SQL error messages are often (perhaps usually?) nonsense, so you have to operate without them.
- An unfortunate amount of downtime, followed by bizarre excuses.
- It's so darn slow, relative to equivalent queries on BigQuery or Snowflake.
- Even submitting a query can take a weird amount of time.
If Databricks-as-an-RDBMS were competing against Teradata, sure, let's have a chat.But we're in 2021, and there's just no comparing the experience of the SQL analyst on Databricks-as-an-RDBMS vs. Snowflake/BigQuery.
I'm excited for the potential of Snowflake's SnowPark (though know little about it). Calling UDFs from SQL means you can create great features for SQL analysts, provided that they can build the momentum to need it.