Great questions. Happy to answer them here.
First of all, this work builds on Hinton et al.’s second paper, the one about EM routing of matrix capsules, from last year: https://ai.google/research/pubs/pub46653 This work is only minimally related to his previous paper (Sabour et al.'s paper) from two years ago!
RESPONSES TO #1:
* The same algorithm also achieves SOTA in another domain, natural language. Same code. I think it’s significant that the same code, without change, produces SOTA in two domains. See the README and tables 3 and 4 in the draft paper. Don't you think this is significant?
* It requires fewer parameters: 272K instead of 310K for Hinton et al. (2018)’s model and 2.7M for the best performing CNN on record (Cireşan et al.); see table 2. That’s 10x fewer parameters than the best performing CNN on record.
* It requires an order of magnitude less training: 50 epochs instead of 300 for Hinton et al. (2018)'s model.
* It’s trained with minimal data augmentation, unlike Hinton et al.’s and Cireşan et al.’s models (the latter, in particular, uses a ton of data augmentation). Also, unlike Hinton’s model, it accepts full-size images instead of 32x32 crops that are 9 times smaller. Finally, we do not measure accuracy as a mean of multiple crops. So, the model has fewer parameters, requires less training, and has greater capacity.
* It seems to be learning a form of "reverse graphics" on its own, from only pixels and labels, without having to optimize explicitly for it. See the README, figure 4, and the 24 plots and captions in supplemental figures 6 and 7. This is rather significant, don't you think?
RESPONSES TO #2:
* As far as I know, the best attempt at recreating Hinton et al.’s work on EM routing is by Ashley Gritzman at IBM, in July of this year -- only a bit over two months ago. As far as I can tell, his model does not come close to matching Hinton’s performance:
https://arxiv.org/abs/1907.00652
https://github.com/IBM/matrix-capsules-with-em-routing
https://medium.com/@ashleygritzman/available-now-open-source...
* There have been a few other efforts, all of which seem to fall short of Hinton's performance. Gritzman does a good job of covering those other efforts in his Medium article. None of these efforts propose any new ideas, as far as I can tell.
RESPONSES TO #3:
* Me too. So does Hinton: https://openreview.net/forum?id=HJWLfGWRb ... and so does everyone else.
* Alas, as Paul Barham and Michal Isard at Google Brain showed earlier this year, currently it can be challenging to scale capsule networks to large datasets and output spaces, in some circumstances, due in part to current software (e.g., PyTorch, TensorFlow) and hardware (e.g., GPUs, TPUs) systems, which are highly optimized for a fairly small set of computational kernels, in a way that is tightly coupled with memory hardware, leading to poor performance on non-standard workloads, including basic operations on capsules. Source: Barham and Isard (2019) - https://dl.acm.org/citation.cfm?id=3321441 (the PDF is available for free download at that link).
* My draft paper mentions Barham and Isard’s work.