The more immediate causal reason is that there is no off the shelf high performance hardware to accelerate these paths. GPUs were made for graphics, not machine learning. They just happened to be really good at running certain kinds of models that can be reduced to a ton of matrix math that GPUs do really really fast.
Something like Intel Xeon Phi or one of the many-many-core ARM designs I've heard talked about for years would be better for more open ended research in this field. You want loads and loads of simple general purpose cores with local RAM. Put like 4096 32-bit ARM or RISC-V cores with super fast local RAM on a die and transactional or DMA-like access to larger main memory, and make it so these chips can be cross linked to form even larger clusters the way nVidia cards can. This is the kind of hardware you'd want.
When I was in college in the early 2000s I played around with genetic algorithms and artificial life simulations on a cluster of IBM Cell Broadband Engine processors at the University of Cincinnati. That CPU was an early hybrid many-core design where you had one PowerPC-based controller and a bunch of simplified specialized cores like a GPU but more designed for general purpose compute. Programming it was hairy but you could get great performance for the time on a lot of things.
Computing is unfortunately full of probably better roads not taken for circumstantial reasons. JavaScript was originally supposed to be a functional language-- a real, cleaner one. There were much, much better OSes around than Unix in the 80s and 90s but they were proprietary or hardware-specific. We are stuck with IPv4 for a long time because IPv6 came too late, and if they'd just added say 1-2 more octets to V4 then V6 would not be necessary. I could go on for a while.
This has led to a "worse is better" view that I think might just be a rationalization for the tyranny of path-dependent effects and lock-in after deployment.