117 karma · joined November 5, 2017
cmd-l cmd-c ctrl-a r <space> u r l : <enter> cmd-t cmd-v ctrl-a h n <enter>
I have set up `r` as keyword to search on Reddit and `hn` to search on Hacker News [1]. `url:` is the Reddit keyword to search in the url field of the items. (In Firefox, you can right click any search field and add a keyword bookmark for it.)Edit: I have FF Studies disabled under about:preferences#privacy. I guess that is the reason why it is not installed on my machine.
That's probably only a problem if it is must faster than everbody else.
> let alone faster than what happens in our environment
That is often not very hard. When a bottle rolls off the table, you can catch it by approximately predicting it's trajectory without computing the precise evolution of the ~10^26 atoms that make up the water bottle. Compression is a corner stone of intelligence. The second corner stone is using compression to choose actions that maximize expected cumulative future reward.
Capsules do require much fewer parameters, they generalize 10-20% better to new viewpoints, they are much more robust to adversarial examples and can better recognize overlapping objects; but, on the other hand, capsules currently require much more training data to achieve the same performance, even though in theory (if they would actually learn inverse graphics) they should require less data, and they add a lot of expensive additional structure (roughly 10X). I am rather pessimistic about whether the approach will lead us anywhere; it seems sub-optimal to model all possible child-parent configurations explicitly. That has a quadratic nature to it and my hunch is that can be done sub-linearily.
I think, it would still be a very "blunt tool" for feature detection. If you are going to compute weighted sums in a convolution anyway (as opposed to just summation in avg pooling or maximum search in max pooling), then the question is really why not simply learn arbitrary feature detectors instead of fixed Gaussian kernels? You can separate Gaussian kernels in x and y direction, which allows you to compute it in 2 * N^2 * K + N^2 instead of N^2 * K^2 operations (with image size N and kernel size K), but in practice, that probably won't give you enough improvement to make up for how few bits of information a Gaussian filter can extract. You would also need to use a very strong sparsity regularizer to get few enough active neurons in the previous layer such that that multiple Gaussians can infer a location. I am not entirely sure it would not work, maybe it is worth a try.
> If you have too many active neurons then, as you say, you encounter aliasing effects, but I think the same is true with capsule networks - they're not expected to handle particularly high-frequency features, are they?
That is a very good point. In neuro lingo, this aliasing is called "crowding". A multi-channel filter kernel (as in standard CNNs) can in principle deal with that by learning filters for representing multiple entities in different spatial configurations within the receptive field, but that requires large amounts of filters and spatial codes which are also not trainable very well in CNNs. Capsules can indeed only represent one entity within their respective receptive fields. I think, capsules fail more gracefully in case of crowding than standard CNNs because the agreement detection can decide on one out of multiple objects being predicted by the capsules below.