If capsules work wonders for you, my first guess would be that you can improve your training of the standard network to make it work equally well.
In general, my hunch is that capsules are still too low level and too much of a local change to make a strong difference.
To give an example, all of the state of the art optical flow AIs are based on building cost volumes and then resolving them. There are edge cases, where one can prove mathematically that reducing the cost volume to a flow direction will make it impossible to produce the correct result. So to make a significant contribution, it doesn't help to use capsules in the feature processing stage, but you need to replace the entire architecture.