It can easily be wildly non deterministic across different cpus or GPUs, or even in the same session, with the same input.
Performance seems to get worse with new releases, and there are frequent subtle breaking changes when using models built with old versions on newer releases.
Tensorflow serving is barely controllable, and requires insane tuning to make it perform the same as pytorch, but provides little to no docs.
The vast majority of models people build just don't work in tensorflow serving either, as you can't reach in with hacky python to mess with internal state.
If you use a custom host instead then you have to deal with literal gigabytes of python dependencies, making your docker images huge.
Memory usage is uncontrollable and causes terrible performance or instant host death. Results vary depending on cpu count, and automatic parallelism can reduce performance.
I just don't understand how Google use tensorflow internally for real world services.