Dedicated accelerators do that already, e.g. google's TPUs, tesla's D1 or apple's neural engine. You must load the data into compute-unit local memory first before executing matmuls. Keeping the weights there and only piping the dynamic data through it saves memory bandwidth.
Something else that might be theoretically possible is -
Large array of FPGAs are apparently used to simulate and verify chips [1], can the same be done to run LLMs? Can we have 0.25 to 1 token per cycle, would the engineering effort be worth it, and would it be financially feasible from a TCO standpoint?
[1] https://www.servethehome.com/amd-vp1902-is-leviathan-fpga-do...
https://www.microsoft.com/en-us/research/project/project-bra...
edit: typo
Not to mention that GPUs already execute in-order (at least any that I’m familiar with). They do have multiple execution pipelines, but instruction fetch/decode is in-order unlike something like a typical modern high performance CPU.
I'm unsure if this would be much of a win.