The problem with the rest of inference is that changes are not trivially correct or incorrect, as they are with the tokenization layer.
BTW throughput is measured for a 12 GiB file. Would be interesting to see the throughput for something more like 32 KiB, with cold start (token cache not yet populated).