I might be missing something but it seems like most of the speed up is from quantization which is commonly used already, and the CPU instance used here isn't that much cheaper (~10-15%?) than a GPU instance that could run the model. For high utilization workloads the extra throughput might be useful though.