The architecture diagram in the article resembles the approach Apple took in the design of their neural engine.
https://www.patentlyapple.com/2021/04/apple-reveals-a-multi-...
Typically these architectures are great for compute. How will it do on scalar tasks with a lot of branching? I doubt well.