Unless you really are doing greenfield development in an isolated application, these considerations often trump any language feature.
Because sometimes what you really need, or at least want, is an overlap of concurrency and other things, and Python is optimal for you for the other things.
If you just need concurrency, sure, but how often is that the only requirement?
Very few people set out going asking themselves about such low level details on day one of a project. Especially something that was an MVP or POC
The answer historically has been c/c++ and bind to python. This work is mainly motivated by one of those libraries wanting to write less c++ bindings and be able to do operations like these parallel directly in python.
A couple issues with multiprocessing is fork is often likely to dead lock especially with cuda and tensorflow. tensorflow sessions aren't even fork safe and while you can fork before making tensorflow it is a restriction to be careful of. It's also comes with heavier cost for smaller tasks that you want to be short and interactive but are still very compute heavy. Relevant quote for that, "Starting a thread takes ~100 µs, while spawning a sub-process takes ~50 ms (50,000 µs) due to Python re-initialization."
Communication/sharing of memory is also more expensive between processes than with threads. One key quote,
"For example, in PyTorch, which uses multiprocessing to parallelize data loading, copying Tensors to inter-process shared-memory is sometimes the most expensive operation in the data loading pipeline."
My experience has been gpus actually compute fast enough that sometimes memory bandwidth becomes a bottleneck and making memory sharing cheaper becomes very relevant. Data transfer costs are pretty noticeable and something to minimize. I've seen a decent amount of interest in zero deserialization formats from this.
edit: Another example if you examine intel's c++ library for ML operations onednn the implementation is not process heavy. It is based on multithreading. I generally see threading as main primitive for parallelism in C++ libraries even though multiprocessing is certainly an option they can take. You will often find single operations (like one matrix multiplication) have implementations that use multiple threads. For individual operations you aren't going to want to use multiple processes to speed them up. That's why tensorflow has both an inter operation thread count and an intra operation thread count when configured.
Specific to your point, recruiting for Elixir talent is a problem compared to more mainstream languages. Recruiting in general is extremely hard at this moment.