Per my understanding of XLA, it provides the right high-level abstraction for compiling the tensorflow computation graph under a range of architectures, CPU, NVIDIA GPU, mobile, etc. but it still delegates to a lower level domain-specific API like CUDA.
This undertaking isn't something to be taken lightly; CuBLAS has some pretty cutting edge architecture-specific optimizations for batching operations for matrix multiplication that came from several years of research - and is arguably a massive competitive advantage of NVIDIA over AMD. Depending on the development state of such an API internally within Google, it could mean that the Cloud TPU isn't going to be ready for wide-spread commercial use for a good while, and is very much still in the research-and-development phase (which could explain why they're only opening it up for the research community right now).