GPUs are massively parallel limited purpose CPUs. You're always waiting for the program counter to get to the end, and getting data to/from Ram.
My project (since the 1980s) is to put compute and memory together, so you can either store 16 bits, or do a 4 bit look up with the same data in every cell (which has 4x16 bits total). Having no memory-compute barrier should allow for pushing data through a program, in a physical sense. Like an FPGA, but easier to program.
Imagine such a chip at a few Ghz, delivering one output per clock cycle for a given task.