Most of what I do is floating point operations, so I'll see >100% performance gains with the avx boost. I will upgrade.
While GPU offloading is extremely powerful in terms of compute, you have to deal with both bandwidth limitations (de facto ~12 GB/s for x16 PCIe 3.0) and the latency of launching compute kernels and waiting for them to complete.
For real time applications this is usually not an option, there latency > some threshold will kill your proposed solution.