Here's a youtube video that explains the new capsule idea: https://www.youtube.com/watch?v=VKoLGnq15RM
(I'll admit that I don't fully understand it yet), but I think the major thing that capsules tries to fix is that a CNN only looks at a small window of the image at a time. Since the capsules aggregate more information, it can learn more general features.
Also, he notes that the paper was done on the MNIST data set (small images), and may not generalize to larger images, but the initial results are promising.