It's beyond me why processor with dataflow architecture is not being used for ML/AI workloads, not even in minority [1]. Native dataflow processor will hands down beats Von Neumann based architecture in term of performance and efficiency for ML/AI workloads, and GPU will be left redundant for graphics processing instead of being the default co-processor or accelerator for ML/AI [2].
[1] Dataflow architecture:
https://en.wikipedia.org/wiki/Dataflow_architecture
[2] The GPU is not always faster: