Debugging a failing test case caused by query running “too fast”
databricks.com
databricks.com
I eventually realised that the routine was fine, but my test data was being generated quickly enough that time.time() (IIRC) returned identical values for all of the dummy records (with the print statements, there was just enough of a delay for there to be a few milliseconds between each one).
I've been consistently impressed with Databricks' approachable blog. This particular post spawned a nice discussion around database design with my son, who has taken a lot of recent interest in all things technological. Keep up the good work.
I keep coming back with every new Spark version to see if the problem has gone away, (wrote it at 2.0.0, so I mean every minor and patch). I looked up what I could online about optimisation in Spark, and applied that.
The business people got tired of us wasting time trying to optimise, and forced us down the lines of SAP HANA and other proprietary marketing hoohah because we need a product that's real-time.
I hope the upcoming version of Spark at least helps reduce latency, perhaps through improvements in the whole-stage code-gen.
A cross join is just two nested loops iterating over one array over another. With 40 cores, each handles 25 billions iterations of the 1 trillion. Assuming each iteration takes 10 CPU cycles, a 6GHZ core can handle 600M iterations/second. 25B / 600M = 41 seconds to run the whole thing.
Yes, 1 second is too fast.
Awesome that they figured it out it's the JVM optimized out the no-side effect computation.
These are also handy to test for SQL injection issues without screwing anything up.
They're also handy to exploit SQL injection issues, in cases where you can't see the output of the query, but can measure how long it takes to execute :)