I think a lot of this can be done now with C++'s template metaprogramming, or in fact using D's CTFE which is quite a pleasure to use. Had Java been not so broken for numeric stuff, its late inlining facilities would have helped too. I have played with this a long time ago and at that time nothing I tried, (not even Fortran) could match it for the same algorithm. Although my fortran would have left a lot to be desired. I had very little idea about what I was doing with the piece of Fortran code.
There is a Fortran name for what I was using, it eludes me now, but the integrator was essentially a coroutine which would yield control to the function computer and resume integrating once it was computed. Its quite amazing that Fortran has coroutine facilities, they call it something else.
EDIT:
Java is not smart about using SSE and such SIMD instructions, this is really a pity. C#, F# are better about this unfortunately its a Microsoft only thing, yes I am aware of Mono. The other problem is that JAVA is over specified. Java compilers have very little room to maneuver. For example arguments of a function are supposed to be evaluated left to right, there goes an opportunity for low level parallelization. C compilers makes no such guarantees and can in principle parallelize such instances. The other problem is that unless one uses low level loops, Java produces just way too many temporaries which on one hand stresses the garbage collector and on the other wastes time instantiating and filling the temporaries only to be garbage collected away. In Java the only form of polymorphism is runtime polymorphism via virtual functions. Compile time polymorphism can be a lot more efficient. For example if accessing (i,j)th entry of matrix is a virtual function that all but ensures that your FLOPS will sink like a stone. Have it as a CRTP in C++ all that will get inlined, and if you are lucky unrolled and then SIMD'ized. That said Java numerics has improved quite a bit lately.
Regarding C's whole program analysis, my comment was about the state of the affairs then, but even now owing to the semantics of the language whole program analysis in C is orders of magnitude more difficult than in Functional languages. The latter has much more semantic information to play with. This was supposed to be mitigated some with the introduction of __restrict__ but it did not quite deliver on the promise. But cat'ing the entire code to one file and compiling it does definitely speed up the runtime, poor man's whole program analysis ! Now you do have support for whole program analysis and link time optimization that this is no longer such a necessity.