https://spark.apache.org/ https://www.iguazio.com/data-science-post-hadoop/
https://spark.apache.org/ https://www.iguazio.com/data-science-post-hadoop/
I've tried building it for my day job, I would rather have a colonoscopy without sedation.
It's more pleasant and dignified.
https://adamdrake.com/command-line-tools-can-be-235x-faster-...
The only real system dependencies are Java8, maven and texlive (and Python/R if you build for that). Then it's `make-distribution.sh` with the appropriate flags. Scala and everything else that is needed is downloaded by maven. The resulting directory is self-contained assuming you have java8 runtime on your target machine.
https://gist.github.com/Mister-Meeseeks/1ebf875b6e1262449cbc...
And good luck running Spark without Hadoop ;)
- Using Parquet files = parquet-mr which is tied to Hadoop MR https://github.com/apache/spark/tree/master/sql/core/src/mai...
- Using S3 instead of HDFS = Hadoop S3a connector
Even if you don't run HDFS and YARN, you aren't escaping Hadoop. And if some configuration goes wrong, and you'll probably need to look into the Hadoop conf files.
The original comment was about the mass of libraries that Hadoop brings in. Spark isn't a solution that allows you to leave the mess. If you try to dockerize spark, you'll still see that you have 300 MB size images full of JARs that came from wherever.
No one is willing to invest the time to untangle the build process and fix compat issues in the software. The project is slowly dying out.