My SQL knowledge is very limited - I had heard of HAVING but not GROUP BY CUBE or COALESCE - but one thing stood out: "The rewritten sql ... ran in a few seconds compared to over half an hour for the original query."
I know there were four million rows in the dataset, but is 30 minutes the kind of run-time you would expect for a query like this?