What you're saying about distributing a computation on a cluster sounds very interesting. I used Mathematica for a hybrid Mathematica/C++ calculation (LibraryLink) where the complexity was handled by Mathematica and the (simple) heavy lifting by C++. I used the standard parallel tools to run it on a cluster, which means that communication was done through MathLink.
I never went above ~70 CPUs, but people say that problems start to appear above that (too many MathLink connection): http://mathematica.stackexchange.com/questions/20356/mathema...
Another possible problem with my solution (LibraryLink, then Mma parallelization) was that it required a Mma license for as many kernels as I was running, even though most of them were only running the C++ code. But that's easy to fix on WRI's side.