Does this mean we have unconsciously developed a language that exposes such relations?
Does this mean we have unconsciously developed a language that exposes such relations?
"It is an open question why our model recovers the concept of sentiment in such a precise, disentangled, interpretable, and manipulable way. It is possible that sentiment as a conditioning feature has strong predictive capability for language modelling. This is likely since sentiment is such an important component of a review."
They go on to frame that as an important consideration for further work like this:
"Our work highlights the sensitivity of learned representations to the data distribution they are trained on. The results make clear that it is unrealistic to expect a model trained on a corpus of books, where the two most common genres are Romance and Fantasy, to learn an encoding which preserves the exact sentiment of a review."
I'm wondering if a "funniness" neuron could be discovered in a model trained on millions of jokes of various funniness, or what sorts of undiscovered meaning there is in other neurons in this model.