Implementing a CNN for Text Classification in Tensorflow
wildml.com
wildml.com
Though I'm curious why you used VALID padding not SAME for the conv layers? It seems like it would be simpler to use SAME.
Also, minor nit: TensorFlow and TensorBoard should both have two letters capitalized
I will fix the capitalization!
Each sentence vector ends up being the length of the vocabulary, so they're already the same length. You can probably drop step #3 in this case.
There is a way to do it without padding, but it's less efficient from a training point of view. You could instantiate a new network for each possible sentence length then share the paramaters between them, and then batch based on your sentence length.
Also, the padding isn't striclty necessary in theory. The feature vector will always end up being the same length, regardless of sentence length, due to the pooling layer. However, Tensorflow forces you to specify the exact size of the pooling operation (you can't just say "pool over the full input"), so you need it if you're using TF.
Is there an example anywhere of how to initilize from the word2vec embeddings?
session.run(W.assign(numpy_word2_vec_matrix)). W would the embedding matrix created in first layer of the code. [1]
Of course you'd first need to load word2vec and filter its vocabulary to match your own vocabulary. That's most of the code and not specific to TensorFlow. You could use gensim [2] for that.
[1] https://www.tensorflow.org/versions/master/api_docs/python/s...
They combine a CNN with a LSTM for question answering on complex, non-factoid questions. Their LSTM+Attention model performs slightly better, but it's a pretty interesting approach.