"NNs will by their nature never be causal (their entire point is that they can approximate everything)".
A NN can sureley result in a traditional causal model.
For example: Build a NN that has the same computational structure as some simple physics law. By giving it training data it then figures out necessary constants. That may converge to the traditional model, which is mirrowed in the network architecture anyway.
So I really can't support saying NN will by their nature NEVER be causal.