So the question becomes, how many other parts of the inference pipeline have left 1000x optimization opportunities lying on the table?
BTW throughput is measured for a 12 GiB file. Would be interesting to see the throughput for something more like 32 KiB, with cold start (token cache not yet populated).