A Deep Learning USB Stick
movidius.com
movidius.com
Movidius makes low power neural network processors for mobile application. The Myriad V1 is used in google tango and the V2 (what the USB stick has) is used in the new DJI Phantom 4.
http://www.theverge.com/2016/3/16/11242578/movidius-myriad-2...
The Myriad chips are interesting because they combine MIPI camera interface lanes on the same chip as a general purpose NN/CV processor and an SDK suite of hardware accelerated computer vision functions (edge detection, Guassian blur, etc).
here's the white paper for the chip: http://uploads.movidius.com/1441734401-Myriad-2-product-brie...
Because programming these chips essentially requires having the hardware, and because the hardware was very hard to come by, programming these chips was mostly limited to Google, DJI, and other big partners.
With this release the everyday developer has access to these vision processing chips, and the barrier to development entry is considerably lower.
This is not meant to replace your titan X gpu.
(Also, of course, this stick doesn't seem to have any kind of connectivity besides the USB to the host computer. How do I connect my camera? Having to shuffle the data from a camera to the stick passing the host computer somewhat defeats the point.)
They have hardware convolutions on 12 SHAVE cores (kind of a DSP core). It means that the chip can run some useful subset of convolutional neural networks very fast and energy efficient.
They also have 2 general purpose SPARC cores, which allows you to have a "normal" program running there. Not sure, how locked the USB stick going to be, and if running your custom program would be an option.
>How do I connect my camera? The chip itself has a couple of MIPI lanes. The USB stick likely does not expose that. And I agree, that's suboptimal.
I think it's just can make sense to push to the market this kind of chip early and don't wait to be bundle with another device. Kudo for Caffe support.
Given the limited availability of phones and products with the Myriad 2, and the fact that you might not be targeting a phone anyway, being able to buy and develop on a cheap USB stick is an incredibly smart move.
If this is for CV, why not use the myriad DSPs from every silicon company ever?
Seems like a buzzword fueled product.
Likewise, I can imagine people using this with an embedded Raspberry Pi.
They arent just for NN, they are general purpose image processing chips with NNs as one option.
http://www.hotchips.org/wp-content/uploads/hc_archives/hc26/...
But not yet, software is not ready for it yet.
The Google Tango uses the Myriad V1. The new tango to be released in may comes with this new V2 chip.
Pros:
- Security
- Control
Cons:
- Resource limitations
- (...)
At 15 inferences per second in fp16 for Googlenet, I'd guesstimate 50-60 GHFLOPs. That would give it very roughly 2x perf/W over TitanX.
It's still pretty interesting, though, since only need to do the training once.
'With Fathom, every robot, big and small, can now have state-of-the-art vision capabilities'
'It means the same level of surprise and delight we saw at the beginning of the smartphone revolution'
'With more than 1 million units of Myriad 2 already ordered'
tl;dr: because, to an outsider, it sure does look like fraud.
http://www.theverge.com/2016/3/16/11242578/movidius-myriad-2...
Seems like the more logical approach would be to have a widget app developers could easily deploy embedded TensorFlow builds in Android & iPhones. Has anyone looked into doing this or found someone already doing this?
Acceleration is needed for training -- not running the models themselves. A quick glimpse of the power used (1 watt) lets you know exactly how much "acceleration" is going on in here. This is meant for tiny devices.
EDIT: My point is that this is a small-run dev board for a chip for some future $19 nannycam. It's not an "accelerator" you install on your PC to put your graphics card to shame running TensorFlow.
EDIT #2: This is another one of those HN threads that's overrun by enthusiasts. Jamming a chip onto a stick is simply how they sell embedded crap now.
Here's a crypto chip that'll really get you guys going: http://www.atmel.com/tools/AT88CK590.aspx
(For example, Google's voice recognition on Android can run when offline.)
This isn't true. Running neural networks (including CNNs) can be computationally and power intensive, and lends itself to the vector operations of GPUs, FPGAs, and ASICs. Putting the computations on devoted hardware could enable embedded applications that simply aren't possible otherwise.
Here's a whitepaper by Microsoft about using FPGA's to speed up CNNs: http://research.microsoft.com/apps/pubs/?id=240715
Article by Google explaining the importance of optimizing neural networks to run on mobile phones: http://googleresearch.blogspot.com/2015/07/how-google-transl...
I see one example in Verilog on github: https://github.com/ziyan/altera-de2-ann/blob/master/src/ann/...
Have you played with Apple Accelerate? They've been baking this stuff into their chips for quite some time. Apple's FFT outperforms FFTW. https://developer.apple.com/library/mac/navigation/#section=...
What happens when 1000 of these are plugged in to 1000 drones and they all communicate? I don't know enough to even guess but perhaps something interesting and dangerous?
Sometimes training is done on the customer site. For example a noise cancellation algorithm may learn audio characteristics of the user's environment to offer better performance.