Decoding the Thought Vector
gabgoh.github.io
gabgoh.github.io
It struck me at the time that the qualities that were expressed most strongly were the ones that ended up having names in our language. But there were others for which I would say to myself, there is something about this group (e.g. those with the greatest expressed value of F124) that I recognize, but can't quite put my finger on.
Of course, I was looking at people through a keyhole, their TV viewing preferences being the only information I had.
Also, I noticed that these "came into focus" most clearly at a certain level of compression (rank).
FWIW
http://stats.stackexchange.com/questions/121162/is-there-any...
The issues being discussed in the essay have been a central issue in some area of psychology and behavioral sciences for some time--how to interpret components such as these.
One thought about your "coming into focus at a certain level of compression" comment: I've done some analyses of these vectors as applied to text samples, and one thing that struck me was how unreplicable some of them were across datasets that should be ostensibly similar (but are not the same). Others, in contrast, reappeared across multiple corpora. To the extent some of these components represent "real" features, they should reappear consistently across different datasets where you'd expect them to. That is, they should be robust to changes in idiosyncratic features of the database.
I thought it was pretty obvious. The atoms are complected (defined - by Rich Hickey of Clojure - as, basically, a semantic that contains multiple interdependent concepts (for example how variables complect state, values, and names)). In fact that's the conceit of the whole idea, the thought vector is being extracted from the sparse matrix, sometimes that sparse matrix isn't that sparse and you will get complected concepts. It was obvious in the earlier pics. One atom might contain a piece of information needed by a hat, that when combined with other atoms makes a hat, but when combined with a different set makes a headband.
For example if you look knives is shared with scissors. It's one atom describing roughly "handheld sharp objects". The airplane atom + many items atom is actually a special mutation for many knives. Where as the many items atom + sharp objects atom is likely scissors. And the sharp object + airplane atoms are probably knife. They are all complected and interdependent.
Sure an atom may generally mean a specific concept but sometimes it will fall back to a combination specific mutation. For example there aren't often many planes, and many planes looks rather like many knives. And there probably aren't ever more than one pair of scissors, and one pair of scissors looks rather like two knives. It's a way to describe 3 things and their number (knives, scissors, plane) using 2 things, an existing counting mutator, and the fact that scissors and planes are often singular. It's a form of semantic compression, quite interesting, and I would imagine domain dependent.
That's my hypothesis anyway.
https://seattlecentral.edu/faculty/baron/Summer%20Courses/AS...
They are "class" modifiers that modify different nouns in different ways.
Also, if all thoughts can be described as vectors and linear combinations of vectors in "thought-space", I wonder what the axis represent and how many dimensions there are. Are all thoughts just a combination of 100 "unit thoughts"?
Really interesting post!
Since the whole appeal of neural networks is that they can model non-linear functions. Why would the autoencoder end up with an encoding that is essentially linear?
So if the NN is well trained, the second to last layer will have linearly separated the different classes (as much as possible, anyway). Earlier layers may not have completely linearly separated their inputs, but they are probably going to lie along simpler manifolds than even earlier inputs.
There's a good blog post on this here: http://colah.github.io/posts/2014-03-NN-Manifolds-Topology/
> Rather curiously, it turns "airplanes" into "knives". I do not understand why this happens.
I would venture to guess that this happens because of the existence of an plural ambiguous "thought" bridging the "airplane" and "knives" concept vectors, along the lines of airplanes -> propellers -> blades -> knives, and the "a group of" vector is causing the system to jump over that semantic ambiguity.