data engineer here, offtopic, but am i the only guy tired of databricks shilling their tools as the end-all, be-all solutions for all things data engineering?
Things databricks offers that makes peoples lives easier:
- Out the box kubernetes with no set up
- Preconfigured spark
Those are genuinely really useful, but then there's all this extra stuff that makes people's lives worse or drives bad practice:
- Everything is a notebook
- Local development is discouraged
- Version pinning of libraries has very ugly/bad support
- Clusters take 5 minutes to load even if you just want to "print('hello world')"
Sigh! I worked at a company that was databricks heavy and an still suffering PTSD. Sorry for the rant.
P.S. here is a simple example of unit testing: https://github.com/alexott/databricks-nutter-repos-demo - I wrote it more than three years ago.