Any chance your workload would see a benefit from porting the code to a GPU?
While GPU offloading is extremely powerful in terms of compute, you have to deal with both bandwidth limitations (de facto ~12 GB/s for x16 PCIe 3.0) and the latency of launching compute kernels and waiting for them to complete.
For real time applications this is usually not an option, there latency > some threshold will kill your proposed solution.