Each sentence vector ends up being the length of the vocabulary, so they're already the same length. You can probably drop step #3 in this case.
Each sentence vector ends up being the length of the vocabulary, so they're already the same length. You can probably drop step #3 in this case.
There is a way to do it without padding, but it's less efficient from a training point of view. You could instantiate a new network for each possible sentence length then share the paramaters between them, and then batch based on your sentence length.
Also, the padding isn't striclty necessary in theory. The feature vector will always end up being the same length, regardless of sentence length, due to the pooling layer. However, Tensorflow forces you to specify the exact size of the pooling operation (you can't just say "pool over the full input"), so you need it if you're using TF.