Curious if anyone has thoughts on going even further: eschewing soft-ware based inference in favor of a purely ASIC approach to a static LLM. Cost benefits? Software level additional, fine-tuneable layers to allow a degree of improvement and flexibility? We are quickly approaching ‘good enough’ for some tasks—at what point does that mean we’re comfortable locking something in for the ~2-4 year lifespan of a device if there _were_ advantages offered by a hyper-specialized chip?