Model serving is trivial, and inference is just memory bandwidth. The cost of serving will be asymptotic to flash read energy.
Having trained on your own chips, that is the impressive part.
Having trained on your own chips, that is the impressive part.