Thread level parallelism is different from GPU parallelism. Different threads can perform completely independent operations at any time. GPU threads must do exactly the same operations, but on different memory locations, at all time. In exchange for this rigidity, we can pack a lot more of them on silicon than CPU. A CPU thread is like a complete individual that can do anything they want. A GPU thread always is part of a pack, and they all move together.
The nice parallelism allowed by Clojure is for CPU threads not GPU threads. It would still need to rely on an external library for tensor operations, for instance ATen [1], the C++ backend of PyTorch.
On the other hand, Functional Programming can be useful to describe the model at a higher level and better handle the scheduling of each component (Convolution, LSTM, etc) on GPU. When training model, the batch size already allows near optimal usage of a GPU cores, however when doing evaluation, this becomes more relevant.