In short, you don't want to spend weeks on hand-optimized kernels, every time you want to implement a new type of layer for a large-scale network, or on an experimental target hardware.
The highest performance approach (on CPU) seems to be JIT assembly (as used in Intel OpenVINO). On Intel, that beats the living Jesus out of everything else, including my work. On ARM it's a free for all, and there's no clear leader, especially in quantized inference. On GPU whatever it is NVIDIA is doing is the right thing to do.
I'm skeptical that XLA/MLIR/TVM-like approaches can come close to, let alone exceed, the performance of hand-tuned kernels, for the same reason why hand-tuned assembly beats the shit of what compiler generates most of the time. I've yet to see it happen in practice. And you just need a few of those kernels, strategically placed where most of the computation happens, as per Pareto principle. For something like TF Google has the resources to get that done. It just chooses not to, to sell you more GPU-hours.
But this will not be a general library right ? You must have only included certain subset of functions of TF or PyTorch or whatever. Autodiffing is also included in certain proprietary libraries like ones from NAG. But I doubt its possible to achieve 50x speedup without compromising on functionality.
It mean you have efficient "crud" operators? This interest me because I'm building a relational language and toy with the idea of tensors as the "table" structure, but get rid of that because updates...
Google certainly has _technical_ capacity to do what PyTorch does, but not the _organizational_ capacity to scrap TF and start over. So they're trying to half-ass it with TF 2.0. It still sucks though.
Perhaps Google could copy PyTorch. But Tensorflow has a lot of overhead, and for both political/technical reasons, there's no easy path to go from Tensorflow to Pytorch.
You could just as easily ask: Why is Google sticking with {Hangouts/(Allo,Duo)/Angular} instead of doing/copying {Zoom/WhatsApp/React}? It's not like Google lacks the technical ability.