Why would you think there's an ASIC/FPGA design that significantly improves over GPUs specifically targeted at running large models already? Where's the win?
The fundamental limit for hardware acceleration are number of gates you can squeeze on a die, right now. (Or, alternatively. memory bandwidth)
analog chips like what MythicAI is developing seem like the next obvious leap forward for deploying inferences broadly. ASIC/FPGA wouldn't be much different than a GPU. ASIC seems like a brittle solution