Text Understanding from Scratch Using Temporal Convolutional Networks
arxiv.org
arxiv.org
No, it is not unique. We have, among other things seen character-level language models (Sutskever et al. 2011) [1] and character-level part-of-speech tagging (Santos et al. 2014) [2]. What is unique are the convolutional aspects.
I am still for from convinced. The baselines are really weak sauce, sure, new datasets and wanting to use the same baseline for all tasks, but a Bag-of-Words model is pretty much the weakest baseline there is for Natural Language Processing tasks. Also, using the 5,000 most frequent words will hurt the BoW model for plenty of tasks since it will cover mostly function words rather than rare nouns due to the Zipfian nature of language. It is pretty much common knowledge that these rare nouns can be far more useful than function words for tasks such as topic classification.
What is interesting to me is that if ConvNet works well both for language and for visual processing that may well be because the human circuitry for processing both are very similar, while formalized grammar is at a different level (like logic) above speech as opposed to the linguistic view of a universal grammar undergirding speech.
This paper looks to just show the major winning aspect of using CovNets as they do not need many features as the deep net learns its own representations of the training data. It more to show CovNets work on more then just vision.
But architeching the pooling layers IS adding complex to the simple input feature set. Therefore the comparison should be of only state of the art ML.
* Compare to RNNs with character level input.
* Compare to dedicated methods of sentiment analysis and topic categorization.