Not sure if you're joking or not? Fib is the classic first parallel benchmark, especially for functional languages. The two tasks in each step are embarrassingly parallel.
> I have trouble believing that any compiler could generate them
Hah! I'd say I'd have a hard time believing compilers can automatically parallelise anything else... because fib is usually the only thing they're shown being able to do!
That baffles me. I think I've hardly read a paper on parallel functional programming that doesn't start with fib.
Here's example that agrees that it's the hello-world of parallelisation.
https://wiki.haskell.org/Haskell_for_multicores
It's trivial for a compiler to automatically parallelise it - for every binary operator evaluate the two operands in parallel - done.
Computing fib with memoisation from 1 will certainly be faster than (parallel) recursion without memoisation and data sharing.
An iterative fibonacci solution runs in linear time. Calculating the N'th fibonacci number requires O(N) serial operations.
A fully parallel, recursive solution without memoization requires that each value in the sequence is computed more than once. Consider the example of fib(4):
fib(4) = fib(3) + fib(2)
fib(3) = fib(2) + fib(1)
fib(2) = fib(1) + fib(0)
You can see, that if we run this in parallel, the value of fib(1) has to be calculated twice. As the tree of operation branches out, more and more duplicate calculations are required.A quick google suggests that the time complexity of the recursive approach is O(2^N).
Absolutely nobody is under the impression that the naive parallel implementation of fib is actually useful code or the most efficient way to do it. You're missing the point if you're suggesting a different way to do it in the first place.
It's just something to use as a running example... like on the original article this whole thread is about.
I think maybe if there hadn't been such a vibe of "what, you don't know X?!?" then it would have just been an interesting fact to mention that parallelization demos often use fib, because it's an easy example to grok (though a confusing one if you already understand the faster method).
Fibonacci computations depend upon prior results, which is why it is a poor fit for parallelization in general. While, yes, it is possible to fork on every recurrence and join to wait for the result, the overhead of that is gigantic and dwarfs any benefits.
Just take a look at the code we’re talking about yourself. See the operator a + b between the two recursions? There are zero data or control dependencies between a and b. They can be perfectly distributed, which is why it’s been used as an example here.
Also read the note directed at the 'algorithms police' here https://cilk.mit.edu/programming/ which I think will pre-empt the point you're going to make next.
It starts two tasks, fib(4) and fib(3). It waits for them to complete. There are now 3 tasks, but 2 are running.
fib(4) starts two tasks, fib(3) (second edition) and fib(2). It waits for them to complete. There are now 5 tasks (fib(5), fib(4), fib(3), fib(3), and fib(2)), and 3 are running (fib(3), fib(3), and fib(2)).
fib(3) (the first one) starts two tasks, fib(2) and fib(1). It waits for them to complete. 7 tasks, 4 are running.
fib(3) (the other one) starts fib(2) and fib(1). It waits for them to complete. 9 tasks, 5 are running.
We keep going like this. Notice how many tasks are not running at any given time because they depend upon other tasks to complete. This is nowhere near an embarrassingly parallel problem.
For an actual embarrassingly parallel problem, consider something like "this 1GB slice of memory contains a continuous sequence of 64-bit floating point numbers. Please square each number in place."
This can be divided among any number of processes trivially, with actually zero coordination.
There's also "nearly embarassingly parallel," where a minor amount of coordination is required in a final step. You can parallelize "compute the max of an array of integers", for example, by splitting it into N chunks, having each one compute their own max, and then taking the max of the maxes.
But Fibonacci is nowhere close to those.
That said, I found a parallel solution that was interesting and should provide a speedup, but with a large increase in complexity of the code.