As far as I can tell, this is almost the same as stacking multiple layers of ensembles, except worse as each ensemble is trained while previous ensembles are learning. This is causing context drift.
To deal with the context drift, Hinton proposes to normalise the output.
This isn't anything new or novel. Expressing "ThIs LoOkS sImIlAr To HoW cOgNiTiOn WoRkS" to make it sound impressive doesn't make it impressive or good by any stretch of the imagination.
Hinton just took something that existed for a long time, made it worse, gave it a different name and wrapped it in a paper under his name.
With every paper I am more convinced that the Laureates don't deserve the award.
Sorry, this "paper" smells from a mile away, and the fact that it is upvoted as much shows that people will upvote anything if they see a pretty name attached.
Edit:
Due to the apparent controversy of my criticism, I can't respond with a reply, so here is my response to the comment below asking what exactly makes this worse.
> As far as I can tell, this is almost the same as stacking multiple layers of ensembles
It isn't new. Ensembling is used and has been used for a long time. All kaggle competitions are won through ensembles and even ensembles of ensembles. It is a well studied field.
> except worse as each ensemble is trained while previous ensembles are learning.
Ensembles exhibit certain properties, but only iff they are trained independently from each other. This is well studied, you can read more about it in Bishop's Pattern recognition book.
> This is causing context drift.
Context drift occurs when a distribution changes over time. This changes the loss landscape which means the global minima change / move.
> To deal with the context drift, Hinton proposes to normalise the output.
So not only is what Hinton built a variation of something that existed already, made it worse by training the models simultaneously, and to handle the fact that it is worse, he adds additional computations to deal with said issue.