* AWS has a managed Spark offering called EMR
* EMR pricing (https://aws.amazon.com/emr/pricing/) is lower than Databricks pricing (https://databricks.com/product/aws-pricing)
* Databricks notebook development experience is better than EMR (but still really basic compared to IntelliJ / PyCharm text editing)
* Both Databricks & EMR have proprietary Spark runtimes
* Databricks is building a Spark runtime in C++ that might be faster (Delta Engine)
* Spark lets you process massive datasets easily, with small teams. 2-3 person teams can build data ingestion pipelines to clean & process terabytes of data a day. It's an incredible technology.
* The difference between the PySpark & Scala APIs confuse the hell out of people
* Whether or not ppl can run Python machine learning models on Spark clusters confuses people
* Overreliance on notebooks causes big issues (no version control, tests, deployment process, dependency management)
The big data ecosystem is constantly evolving and you need to study constantly to keep up.