The tail calls in question are C tail calls inside the inner interpreter loop. They have nothing to do with Python function calls.
If your benchmark setup is easy to re-run, it would be awesome to see numbers that compare the tail call interpreter to the build where it is disabled, to isolate how much improvement is due to that.