Furthermore, I think the results from GPT-2 and similar language models show that researchers have found a scalable technique for sequence understanding. They are likely to just work better and better as you throw more data and training time at them. Imagine what GPT-2 could do if trained on 1000x more data and had 1000x more parameters. It would probably show deep understanding in a great variety of ideas and if prompted properly would probably pass a lot of Turing tests. There is evidence that this type of model learns somewhat generally, that is, structures it learns in one domain do help it learn faster in other domains. I am not sure exactly what would be possible with such a model, but I suspect it would be extremely impressive and meaningful economically.
I think we are likely to see that type of progress in the next year or two, and for there to be no AI winter.