Parallelization would get you a factor 1/k (k being the number of cores) in favor of the classical algorithm at best.
Exactly. The question is what would be the actual real life speedup. If it is not easily parallelizable then it becomes much more interesting finding.
I am not familiar with these algorithms, but name of QMC would suggest that this is an embarrassingly parallelizable problem, so my interest might be just from my ignorance.
The real issue is whether if there is an exponential difference between quantum annealing and classical algorithms
Yes, I know and that was not the focus of my comment. Sorry.