"Wife" sounds exactly the same in both places, so all this did was copy the exact waveform from one point to another. Nothing is being synthesized. If this is all this app can do, it would be quicker and easier to do this manually.
That little "guh" noise at the beginning of the first "wife" could also be manually cleaned up and pitch/formant shifted to sound more natural with respect to its position in the sentence.
https://www.youtube.com/watch?v=I3l4XLZ59iw&t=3m54s
The word "Jordan" is not being synthesized. He was recorded saying "Jordan" beforehand for this insertion demo and they're trying to play it off as though this was synthesized on the fly. This is all a scripted performance. Jordan is phoning in his task of feigning surprise.
https://www.youtube.com/watch?v=I3l4XLZ59iw&t=4m40s
Here they double down on their lie. The phrase "three times" was clearly prerecorded.
If Adobe wanted to put the bare minimum effort into trying to convince anyone this was a real product that exists and is capable of synthesizing speech on the fly, then they'd toss a beachball around the audience and have them shout out words to type.
This is a fraudulent demonstration of a nonexistent product and an audacious insult to everyone's intelligence. Adobe is falsely taking credit and getting endless free publicity for a breakthrough they had no hand in. They are stealing the hype recently generated by Google WaveNet and praying they'll have a real product ready by whatever deadline they've set for themselves.