Also if you haven't looked into the properties of how exactly a RNN Transducer functions, I highly recommend doing so. They help resolve a great deal of problems that traditional RNNs and CNNs are unable to deal with.
The RNN-T is a nice idea though, if I understand it correctly it's another approach to the alignment problem. In CTC, you are generating sequence like TTTTHHHEE CCCAAATT, which mean that your language model must deal with these repetitions, and you can't train using text without repetitions. In RNN-T you are learning to advance the cursor on either audio sequence or text sequence so as to maintain alignment, kind of like you do when you merge-sort two sorted lists, therefore it outputs THE CAT, and you can use a standard language model.
Hoping paper details suffice and help with the parameter search, and happy to respond to questions over e-mail. Would love to see an open-source implementation with local or directed attention built out!
Gated convolutions as LM is similar to RNN-T idea [1], but you have to deal with softmax, so I'm not sure how well this would work in practice, especially on a mobile processor.
The primary reason to be interested in convnets for speech is computational parallelism, not because they have especially strong results for accuracy.
I work in the field, a more accurate summary would be that there are a number of viable architectures that currently get fairly similar accuracy, but that have other pros/cons with respect to streaming, memory use, parallelism, model size, integration with external language models and context, complexity of the decoder, friendliness to different types of hardware etc.
- No comparison is given of number of model parameters. If optimizing strictly for model size, RNNs tend to be nice and compact.
- The computational advantage of the CNN at training time is throughput. The advantage of RNN at decoding time is streaming latency. Running the CNN frame by frame as they are received removes the ability to run frames in parallel and if the CNN is larger, it will run slower, and depending on its receptive fields it may not even stream well at all.
- That particular CNN system uses a strictly external LM that is not jointly trained and has an additional hyper parameter at decoding time to weight the LM that requires additional tuning.
- It is still autoregressive in the beam search, so the LM will still be run many times sequentially adding tokens just like an RNN LM, and is likely to be more expensive. The throughput advantage a conv lm has in scoring whole sentences is totally lost. In fact, there doesn't seem to be anything special about the choice of a conv lm for that paper except that it is fun to make all the parts convolutional.
- CNNs frequently require more total flops, but are high throughput on eg a GPU because they expose so much parallelism. On an embedded CPU this can be a bad tradeoff.
As a side note, there's no reason that CNN architecture, which in the paper is trained with a close relative of CTC and is decoded identically to a RNN CTC AM + external LM, couldn't be trained as an RNN transducer. Despite the name neither the am nor lm have to actually be RNNs.
Recurrence can help with robustness in some other very important ways as well.
Citations for this dates from the 80s and 90s. I don't know the best reference offhand. You could look at some old Hinton stuff if you're a fan. Lots published on this.
Nothing like this has been published AFAIK.
After you have the results of this experiment you can try to explain them with attractors and what not, but I would be surprised if there was much difference. Would make a good paper though!
Unless you're dealing with opinions, I disagree. The onus is on the person trying to give evidence to actually give evidence.
const mapTrans = fn => function* (x) {
yield fn(x);
};
const filterTrans = predicate => function* (x) {
if (predicate(x)) {
yield x;
}
};
const dupeTrans = n => function* (x) {
for (let i = 0; i < n; i++) {
yield x;
}
};
Clojure just observed that an isomorphism of these functions under Church encoding looked like f<B> -> f<A> for a special parametric type f, so could be composed with ordinary function composition, albeit backwards.An RNN transducer is fundamentally three functions:
One takes a list of recent As to some C1.
Another takes a list of recent Bs to some C2.
A final one takes a C1 and C2 to a B.
Rather than mapping over each input independently the RNN-T is allowed to learn something about the relationship of recent outputs to the next output, and the relationship of nearby inputs. Clojure transducers thus have an order insensitivity that RNN transducers are allowed to be sensitive to.