Google AI hadn't published any major papers on ASR or TTS since WaveNet, even the new WaveNet google demo'd a few months ago weren't even close to this demo's voices. The use of "Uh Huh" "Hmmm" are called back channeling, it's moon shot away given the current state of technology. It requires very low latency and precise timing so it doesn't cut the speaker off.
I believe if this demo were real, it was tested and cherry picked from a very particular environment