You could specify computation in high level API like TensorFlow, and then have the framework pick the best implementations available (ie, CuDNN for GPU, MKL for CPU, something custom for TPU)
"something custom for TPU" is what seems interesting here. Researchers working on the TF research cloud probably wouldn't learn much about the sorts of workloads this architecture is suited to, if all they have access to is the high-level API. Would whatever iteration of the TPU API they have internally be something they'd release soon, at least to partners?