After you shard some of the computation out to different processors to filter before single core aggregation, then you can push down some pre aggregation to limit the communication between parallel threads, and it keeps going!
Super interesting to see parallelism for reading logs this way, like for when you have too many logs for one core and too few for Spark