Recurrent Neural Networks Hardware Implementation on FPGA
arxiv.org
arxiv.org
I'm left curious on the performance gain factor when scaling the network in terms of layers and units. Would the performance gap widen as the RNN grows?
They do say in the text that "Figure 8 shows the expected speed up, assuming the data throughput is high enough to handle the parallel processing" so take it with a grain of salt. There could and most likely will be (as there always is) other factors that prevent ideal scaling.
Any algorithms working with images need FPGA implementations to be quick enough, and there are a lot of convnets in use there.
There's a chapter on convnets here: http://www.cambridge.org/us/academic/subjects/computer-scien...
RNNs are nothing special in particular, at least these with small number of layers and nodes.
HMMs or CRFs are easier to handle for sequences, and would probably work well, and there are FPGAs all over the place of these models.
Also, my impression was that FPGAs are much slower than GPUs for neural nets. Unless you're talking about really high end chips like Stratix 10 from Altera, which cost over $30k. Power consumption is a different matter though.
They won't tell you it's NNs but it is. Sony distributes a lot of movies, they use their database of movies to train the upscaling models (which is obviously NNs https://github.com/nagadomi/waifu2x ) and then put the chip in the TV.
It's something almost equivalent to storing Pride and Prejudice and Zombies in your TV in 4K, and then reproducing it when they match it with what's playing on TV.
I don't quite understand what you mean by: "It's something almost equivalent to storing Pride and Prejudice and Zombies in your TV in 4K, and then reproducing it when they match it with what's playing on TV."
Can you explain? What exactly do they have to store in the TV?
Imagine they stored all of the movies they distribute in 4K in your TV. Whenever the movie is displayed they just find it in the database and reproduce the 4K version.
Of course, that is very much infeasible, it requires too large of a storage.
What they do is something like Waifu2x I linked, they train a neural network to learn how to properly upscale movies. They put the "algorithm" on the chip, and its fast enough for real time usage.
When you see side-by-side for these upscalers, it's obvious they aren't that good in deterministic upscaling.