The threat to Nvidia from Intel is in the Nervana chips they recently acquired. Those are presumably using HBM2 and could potentially beat GPUs for neural net training performance.
The threat to Nvidia from Intel is in the Nervana chips they recently acquired. Those are presumably using HBM2 and could potentially beat GPUs for neural net training performance.
I'd recommend reading this paper about matrix math on GPUs from friends of mine way back in the day: https://graphics.stanford.edu/papers/gpumatrixmult/gpumatrix... . While NVIDIA has built much larger register files and L2 caches since then, a modern Xeon still is unbeatable when something fits in L2 or even L3 cache.
I believe streaming the input data in to update the weights is usually done with non Cache polluting instructions (mm_stream equivalents with the NT hint), so it's not hard to keep it fed.
Trying to peg this onto one of their CPU's will have a result much like they have had in pegging the graphics processing stuff onto their CPU's. That is, very meagre results that is fitting for only the lowest resource intensive things that a consumer may want to do.
30MB to hold a model in is almost nothing. Even 12GB is insufficient to train something like imagenet.
But in the same vein as their iGPU implementation, it may be just enough to satisfy the requirements of the average consumer, and still be able to make a large dent in Nvidia's market share.
But while FP16 is useful for audio and imaging, FP16 nearly killed NVIDIA over a decade ago* when it lacked the dynamic range for DirectX 9 HDR effects in contemporary games without banding. FP32 was more than enough for the task, but power hungry, and thus NV30 could be used to figuratively fry eggs while AMD GPUs had FP24, which was just enough for these effects.
These days, it's all going in the opposite direction w/r to deep learning, but I find it ironic that INT8/INT16 is missing from the Tesla flagship P100, but present on GP102 and GP104, the consumer GPUs (and yes, I know about Tesla P40, but that lacks fast FP16).
I agree with the top poster that Intel could make quite a comeback here given how far it's currently behind. I also agree that it would be hard to dethrone NVIDIA without higher bandwidth memory, but I don't think that's necessary, a bloody nose is more than enough to turn heads IMO.
Also backpropagation makes you iterate through the whole model at each step
If future Xeons integrated RAM like Knights Landing does, this would make GPUs less interesting. But KNL can already be used as the main CPU, if I remember correctly.
Intel's offering costs as high as €4000 per processor, which come with a meager 8 cores.
NVidia's offering sells, right now, for less than €1000 a pop.
Furthermore, the performance of each of NVidia's GPUs falls somewhere between 5 and 20 teraFLOPs. A Haswell Xeon gets you about half a teraFLOP for around 1/3 of the power consumption of a NVidia Tesla P100 GPU.
Comparing GPUs with Xeons makes no sense.
On the Nervana front, outrunning a GPU for neural nets is not that hard with an ASIC. I checked your profile, I bet your employer knows something about that :)