I work as a computational mathematician. Part of the problem on my end is that while the libraries have gotten better, they have not contained the operations that I need for good algorithms in my field. As an example, when GPUs first came onto the field, cublas computed a subset of the BLAS routines, which are pretty fundamental for putting together known, good algorithms. Eventually, that got better, but then the operations needed to fit onto a single GPU. Eventually, cublas-xt came onto the market and that allowed parallelization across GPUs, which was great. However, it still didn't have all of the needed matrix operations. At the moment, I'm not sure where it currently is.
Of course, matrix algebra was only one of the problems. BLAS supports the writing of LAPACK routines, which tend to deal more with dense factorizations and eigenvalue problems. I believe more factorizations have been implemented recently, but I'm not currently up to date. Nevertheless, that was absolutely a bottle neck for the longest time. Yes, iterative methods don't need factorizations, but a good fraction of the preconditioners do.
Then, of course, this speaks nothing of the sparse linear algebra problem. There's multiple ways to do it, but many of the good sparse factorization routines need dense factorization routines, so these needed to come onto the market first and that took time.
And, to be clear, I know that there are multiple, good teams working on this. It takes time. And, there's a huge number of operations. If you're bored, go look at the manual for Intel's MKL and see the number and variety of operations that it provides. Those operations are there because people like me need them to do our job. I'll also agree that the kinds of operations we use to do our job will evolve over time with hardware. However, matrix algebra, factorizations, and eigenvalue problems lie at the very core of applied mathematics and expecting the mathematics to rework the last several hundred years of practice to accommodate the lack of tooling from the GPU manufacturers isn't realistic either.
Anyway, if someone knows the current state of what's possible, I'd love to hear. Selfishly, what it really boils down to is what operations (algebra, factorization, or eigenvalue), how big (how much memory or on multiple GPUs), and dense or sparse.