Spark is sort of dead though. Dask looks to be the way of the future. In part because doesn't take a zillion parameters to tune and consume a bucket of resources just for overheads. Good luck.
The venn intersection of conditions where spark makes sense is really rather narrow. A single high spec instance running leaner tooling will generally meet one's requirements while blowing spark out of the water in terms of perf and cost.
Operationally, spark is a huge PITA, hence databricks and a host of other offerings, I guess including this one, to try to manage the pain. Meanwhile something like dask-kubernetes will cater to the same use case with significantly lower operational complexity and again much higher perf and cost efficiency.
I can't really think of a scenario where I'd choose to use spark on a greenfield project today.