Programming languages (at least in the context being discussed) are just a means to model a problem domain. Some people find it easier to model domains in strongly typed functional languages but at a much larger percentage do not. Let's for argument's sake say it's 1% of all programmers. That is still a huge number of people though.
In other words languages like Haskell, Lisp etc. are precisely good for modeling the problem and its solution, so it can make sense to use them for that even if later development is best done translated into a simpler-to-use language.
To be fair, purportedly "easy" languages such as Java aren't actually easy. True, a lot of people can write (or copy paste from online articles) crappy Java code full of bugs and performance issues and nobody will notice since this low quality seems to be some kind of industry standard. But there aren't that many good Java developers either.
A similar argument used to be made about PHP: anyone can use it. To write poor software, that is.
As for Haskell: it's different enough to Java-like and C-like languages that it's difficult to wrap your mind about it if you come from those languages. But if you start fresh, it's not that hard. Some years ago someone here on HN posted pretty good results when teaching Haskell to kids. Young kids pick it up surprisingly fast. It's not inherently difficult, and in fact some of its concepts may be easier to grasp than traditional imperative languages. (And no, you don't need to understand category theory at all in order to program with Haskell).
I'm going to respectfully disagree with that.
I do understand the point you're making. In fact, my own experience of having taught beginners without preconceptions to program is consistent with that point.
However, the "scale" of a language (including all the strange edge cases) is very relevant for production use and particularly maintenance. Even if you stick to a core of widely understood language features and libraries yourself, you might still have to maintain code written by someone else that uses more obscure features or complicated interactions between features.
For Haskell, that "scale" is usually determined in practice by GHC plus whatever language extensions anyone has turned on plus whatever libraries they've found from Hackage or elsewhere. Those language extensions are numerous, and they can not only add new language features but also change existing code to behave in different and potentially unexpected ways. As a platform for research and development of programming language concepts, that is a benefit. As a practical programming language for production use, it is a huge liability.
But yeah, I agree, comparing the haskell you learn at your first programming course, and the production haskell, it's like 2 different worlds
In the case of Haskell, I'm not sure whether that's a distinction without a difference.
How many people, offhand, know the difference between a Runnable, Callable, Function, Consumer, and Supplier? I include all of them because literally the only meaningful difference between them is whether they take an argument or not (Function and Consumer do), and whether they return an argument or not (Callable, Function, and Supplier do).
How many people know the difference between a Future, a CompletableFuture, a RunnableFuture, a ScheduledFuture, a ScheduledRunnableFuture, and a FutureTask? I've done async code in Java; I still can't keep them straight.
Future is something that will eventually produce a value.
CompletableFuture is a Future that you can register a callback on for when it completes.
RunnableFuture I'm not completely sure on, but I'd guess it's a Future that will complete but won't produce a value (basically a Future<Void>).
ScheduledFuture is a Future that you can specify when it should run in the future.
ScheduleRunnableFuture is a RunnableFuture that has a defined time it should start running.
FutureTask is a RunnableFuture you can cancel. OK, so I cheated on this one, I didn't actually know what it was and had to look it up.
Edit: I looked up RunnableFuture and I was half right. It's something that's both a Runnable and a Future, so it has a run() method that causes the future to be evaluated but doesn't return its result at that time. It actually seems pretty niche and probably only useful if you already have some API that's expecting a Runnable, otherwise it's exactly equivalent to calling get() on a Future then ignoring the returned value.
Also, as a note, FutureTask is the only implemented class out of those; the others are all interfaces.
Java is complex as well, it has a 800 pages manual, and that's a language only and no libraries. It has a complex memory model, quirks like pointer-only parametricity, quirks to bypass quirks like autopromotion to help with pointer only parametricity.
Oh, and java have pretty much obscure extensions as well. Ever heard of metaprogramming via Java Reflection?
Haskell without extensions is 300 pages only and way more simple. And extensions are not that numerous and often intersecting.
Simple Made Easy by Rich Hickey
"and a lot of programmers don’t want to learn it" - also opinion, but I'd be very curious to also know how the quality of those who do want to learn it compare to those who don't.
And it's a lot of thousands - don't merely count the cost of training. Count the salary they are being paid while they learn.