The benchmarks are from GCP, where the vTPM is implemented in the hypervisor rather than on something that's plausibly an 8051[1]. Doing this on actual client hardware is going to be a bunch slower.
[1] Typically ARM these days, but most system vendors aren't picking TPM vendors based on performance