It is my understanding that just baking the model itself into silicon only gives moderate gains because memory bandwidth remains a bottleneck.
Think it's called Askjimmy or similar.
(Not that I believe it, it writes too well for GPT-3.)
Hosted frontier models from two years ago would be much faster today, too.
Someday, I imagine model weights could even be encoded as analog resistors (memristors or similar) for even greater density
The trick is that every compute element in their system has it's own small pool of ROM, instead of putting all the ram behind a common pipe. ROM is just used because it's the densest kind of memory that can be fabricated on the same process as their logic.