I am generally not critical of people posting benchmarks. Neither am I trying to rain down on your parade but seriously you do not use Apache Spark to process 3.8G or 15.6G of data. It is too trivial.
We have to start somewhere and if you think we should compare against something else please tell us what you think we should compare ourselves to! We are already looking into some of the other tools that have been mentioned here in the comments.