Hackers, and I'll assume all those vaguely tracing their lineage from MIT, have never been associated with FP. There has always been a fascination with Lisp, but that's the extent of it. In practice, most of their software legacy has been in firmly imperative and procedural languages, but with touches of a Lisp here and there.
It's precisely CS and academia where functional programming has been thriving. John Backus' research on function-level programming, Robin Milner on ML, David Turner on Miranda from which Haskell is almost a direct successor, etc.
So on the whole, I have the exact opposite impression that you do.
If you learn type theory and semantics you also learn functional programming (and will probably like it, even if you don't use it much in practice).
If all the world is a Java VM to you, then... well, you might have a different view of the world.
Using a function style or a higher-level language sometimes makes it easier to modify the program (and hence use a better algorithm).
I remember people swearing by assembly language back in the 80's, thinking C would make their programs slow and bloated. In reality, it was usually the other way around because C code is much more malleable than assembly. You might want to give FP a second look.
https://en.wikipedia.org/wiki/C%2B%2B_AMP
However, my best algorithms are stylistically equivalent to a map-reduce with combiners entirely in GPU-space. The map tasks themselves carry all the scope I need. They are independent units of work executed by independent warps (synchronized groups of 32 GPU threads) whose results need to be reduced to floating point numbers. That reduction is done with 64-bit fixed point atomic ops (my combiners) because:
1) This insures the sum is deterministic
2) Ever since Kepler (GK104), it's much faster than reduction buffers and uses a fraction of the memory
3) It's possible because I can precompute the expected dynamic range and adjust the fixed point exponent accordingly
Now given that, do you see a functional programming equivalent that significantly improves on this design? I haven't yet.
Who cares if you implement the combining with atomic ops or by sending data to a process that does the combining or whether you stuff the data in a buffer and then reduce it afterwards?
(Of course you might care for performance reasons -- but conceptually it's the same thing.)
1) "Who cares" is not the sort of thing you want to say to someone who cares about performance because:
2) Sending the data to buffers for subsequent reduction was the first implementation (2009). But it was a memory hog and a 5% or so slowdown to perform the reduction subsequently rather than concurrently with Atomic Ops and 64-bit Fixed Point (2012).
But I guess what we're arriving at is that this is essentially a functional design using imperative code? I can live with that.
Yep. That's precisely what it is.
I think this preference has more to do with the individual's past experiences and future goals than it does with his/her level of formal CS education.
[1] Although even this varies greatly from school to school, the school where I got my MS used OCaml in the undergrad algorithms course for awhile. My undergrad algorithms course used C, so again no OOP.
Functional languages seems to be one of the hardest things for people without a CS background to pickup in my experience.
I really like functional programming, especially in Standard ML and Haskell. I didn't learn ML until I'd been programming for about 15 years.
Edit: beadder grammer.
[1] At least the Haskell/ML kind.