Using a function style or a higher-level language sometimes makes it easier to modify the program (and hence use a better algorithm).
I remember people swearing by assembly language back in the 80's, thinking C would make their programs slow and bloated. In reality, it was usually the other way around because C code is much more malleable than assembly. You might want to give FP a second look.
https://en.wikipedia.org/wiki/C%2B%2B_AMP
However, my best algorithms are stylistically equivalent to a map-reduce with combiners entirely in GPU-space. The map tasks themselves carry all the scope I need. They are independent units of work executed by independent warps (synchronized groups of 32 GPU threads) whose results need to be reduced to floating point numbers. That reduction is done with 64-bit fixed point atomic ops (my combiners) because:
1) This insures the sum is deterministic
2) Ever since Kepler (GK104), it's much faster than reduction buffers and uses a fraction of the memory
3) It's possible because I can precompute the expected dynamic range and adjust the fixed point exponent accordingly
Now given that, do you see a functional programming equivalent that significantly improves on this design? I haven't yet.
Who cares if you implement the combining with atomic ops or by sending data to a process that does the combining or whether you stuff the data in a buffer and then reduce it afterwards?
(Of course you might care for performance reasons -- but conceptually it's the same thing.)
1) "Who cares" is not the sort of thing you want to say to someone who cares about performance because:
2) Sending the data to buffers for subsequent reduction was the first implementation (2009). But it was a memory hog and a 5% or so slowdown to perform the reduction subsequently rather than concurrently with Atomic Ops and 64-bit Fixed Point (2012).
But I guess what we're arriving at is that this is essentially a functional design using imperative code? I can live with that.
Yep. That's precisely what it is.