Intel Announces Movidius Myriad X VPU, Featuring ‘Neural Compute Engine’
anandtech.com
anandtech.com
http://eyesofthings.eu/wp-content/uploads/deliverables/EoT_D...
I am greatly amused at the thought of 2160p60 MJPEG. That must be extremely efficient…
I wrote a multi stream mjpeg decoder some years ago and it could run three MJPEG streams at 1920x1080@30 on a 2008 Macbook Pro (Core 2 Duo, 2.5ghz ).
And, of course, it's much easier to build a desktop with far more than four CPU cores these days.
The encoding side is very expensive with H.264, but as I understand it a lot of work goes into building the right reference frames so that higher compression and faster decode can be achieved.
Given the target market for this chip, it's a perfectly logical and highly useful feature.
Ive been bitten by a few of their "Maker things". Faulty firmware, no support, no documentation, faulty I2C.. All sorts of things. And unless I have a timeframe how long they plan to support it and sell it, I just can't justify using it in any product I have or plan to build.
Which makes me sad, because I was Intel blue-badge for 11 years back in the day. But the current management seems to be sprinting aimlessly after mirages.
I wouldn't say they're sprinting for mirages. They're hastily entering solid markets but relying solely on their name and prestige. They're not making any investments, beyond marketing, and when the customers don't materialize overnight they're exiting hastily rather than figuring out what they're doing wrong.
Bigger companies have a bad habit of comparing the early results of experimental internal businesses with the ROI of their mainline business and cut the cord before the product/market was actually ready to be flooded with cash to scale up.
It's funny how the business classic "Innovator's Dilemma" was written largely about companies like Intel who exist in traditional technology markets with predictable evolutionary modes... yet they still suffer from the same lack of 'intrapreneurship' mentality by treating internal startups with immature markets as if they were mature product lines.
They should probably stick to acquisitions of real startups in the growth stage or actually stick-it-out for the long run with the markets they invest in, rather than looking for high growth opportunities within a short-timeframe or nothing.
Burning early adopters is never a wise choice if that market turns out to have legs.
You are right that this is showing classic signs of not giving innovative business units enough runway to find their product-market fit. That's really hard to do given Intel's culture.
Craig Barrett did a lot of damage. One Intel mid-level manager described him as a pinata. "Who ever hits him the hardest gets the most candy." So managers reported synthetic disasters in order to get more funding. Paul Otellini was great, I have huge respect for him, but his tenure was too short lived. Otellini understood how to build a market.
The farther a company gets from it's early roots the harder it gets for them to recreate the 'early days'. Unless they get a shake up in management. But the typical people good at startup culture would get killed pretty quickly in bigco corporate culture.
Altera-- Not actually a bad decision, but it seems to have done nothing but bolstered Xilinx's position.
MobileEye-- A great way to burn ~$15 billion. The fact that Intel felt that it had to purchase that company with really no defensible advantage other than its maps highlights the huge challenges facing Intel if they're unable to cultivate an internal SDC team.
Movidius-- Legitimately I have no idea why they paid so much for a glorified DSP. But then again, so are most "Deep Learning Processors" right now.
Curious, do you know alternative "DSPs" able to achieve similar results for Deep-Learning, Computer Vision or Machine-Learning algorithm acceleration?
I would highly appreciate a real answer, because I intended to buy one of these movidius "sticks". And further accelerate the Laptop with an eGPU.
That being said, I think pretty much everything I say (admittedly with some hyperbole) is backed up by data.
So I would recommend at the moment, for a shipping processor, get a GPU.
A few notes though:
Movidius is inference only. That might be useful if you have the specific requirements that needs low power, high speed inference and also somehow has to have a x86 CPU.
If you want high speed, low power inference and don't need x86 then the NVidia Jetson wipes the floor with it.
If you want low speed (~2 inference/second), low power then RPi is a good option.
If you want high speed, you need GPU(s).
Granted, it's 2x the price of a Rpi 3. All together that's about $100 USD. And NVidia just announced the TX1 SE Devkit at $200 USD. I have a TX2, but the TX1 will definitely do better at a higher power/size profile.
The MCS only supports Caffe as well, while the TX1/TX2 will support a wider array of DL frameworks (as well as FP16 support since it's a Tegra).
Movidius have promised TF support I think.
So in summary:
⇒ low power, high speed inference on x86 CPU 🡺 Intel Movidius
⇒ low power, high speed inference non x86 🡺 NV Jetson
⇒ low power, low speed inference 🡺 RPi and similar
⇒ high power, high speed inference 🡺 GPU(s)
NVidia Jetson is the obvious winner, if there is no viable alternative, however Jetson is quite expensive and therefore I can't go that route. I want to speed up training & inference. eGPUs are more or less affordable, but low power training & inference at medium or low-cost would strike me as a clear winner.
Sorry for the late answer, nonetheless, even though I didn't get an alternative DSP that offers similar advantages as Movidius, I'm still grateful for your insightful comment.
HOWEVER, If you are careful, for some models you can get cost benefits by training on cloud CPUs. See http://minimaxir.com/2017/07/cpu-or-gpu/
For the most part when you have hardware that supports one of the binary standards say FP32 you'll have a compliance sheet, more often than not it will not be 100% compliant for example NVIDIA GPUs were not IEEE 754 compliant when Tesla came out, 2nd Generation Tesla cards were FMA compliant, div and sqrt operations were not IEEE 754 compliant until Fermi.
So the GP was correct, back in the old DX9 days when AMD went with FP24 and NVIDIA with FP16/32 both implementations were proprietary, and this is still the case, they just often offer a 754 compliant mode. You can disable 754 compliance to run faster (sometimes considerably so) e.g. in CUDA you can use the NVCC flag --use_fast_math to do so.
But yes, I agree, that criticism is stupid. In fact, I would argue that if you use the IEEE standard for a deep learning processor, you're the one that's stupid.