PS: I do not use C beyond reading some of its code for inspiration, so kinda unaware
A modern C optimizing compiler is going to go absolutely HAM on this sort of code snippet if you tell it to and do next to nothing if you don't - that's, roughly, the explanation for this little discrepancy.
Edit for Twirrim: on this system (Ryzen 7, gcc 11): "-O3": 350ms; "-O3 -march=native": 208ms; "-O2": 998ms; "-O2 -march=native": 1040ms.
Edit 2: Interestingly, changing the C from float to double produces a 3.5x speedup, taking the time elapsed (with "-O3 -march=native") to 58ms, or about 12x faster than JS. This also makes what it's computing closer to the JavaScript version.
(But to be sure, I just ran it again with an output and got the same value.)
No flags: 1843ms
-march=native: 2183 ms
-O2: 423 ms
-O2 -march=native: 250 ms
-O3: 425 ms
-O3 -march=native: 255 ms
O3 doesn't seem to be helping in my case.
No. Float is half the size of double.
Haven't looked closely at the code or tried it, but with -O3, -fopenmp and a well-placed pragma the performance could increase many-fold.
Heck, with NVC++ you could offload that thing to a GPU with minimal effort and have it flying at the memory bandwidth limit.