It says right at the top: "we get 91.8% accuracy versus the previous best of 90.2%" on a standard sentiment corpus. In addition, their method needs less training data than previous approaches.
* Were there any insights learned that give a deeper understanding of these phenomena?
The main appeal lies in the fact that a model trained on a (1) different and (2) very general task basically "in passing" also learned to predict sentiment (i.e., a specialized task that more or less arose from the domain the general model was trained on), and pretty much through a single neuron (out of the 4096 used). The authors speculate that this might be a general effect that could also be transferred to other prediction tasks.