And there are still plenty of cases where this will leave performance improvements of an order of magnitude on the table, because the compiler isn't smart enough to work out the optimal sequencing of operations. Often the computational core which benefits from running on a GPU is small and its complexity low enough that, if performance is a priority, you are better off writing it in an imperative low-level language as a library which can be accessed from higher level code.