I don't know if they sunk all their energy in the new fabs or not ..
I don't know if they sunk all their energy in the new fabs or not ..
Intel has AVX-512, which is in my opinion one of the greatest innvoation in the x96 instruction set for a long time. The problem is that Intel wants to introduce AVX-512 only very slowly into consumer CPUs (slowly beginning with Cannon Lake, which will only be available in small quantities for some time if rumors are to be believed). Even more: For some parts of AVX-512 (e.g. 4FMAPS and 4FMAPS which are very useful for deep learning), Intel seems to be willing only for special expensive accelerator cards (Knights Mill) to include them to segment markets even further.
Similarly it is often complained that Intel offers no ECC support for the Core i... series of CPU (IMHO this complaint is rightful). Again: Intel does offer it, but only in the much more expensive server/workstation CPUs (Xeon).
So in my opinion Intel is sitting on a lot of options and brains - but they seem only willing to sell this in the really expensive CPUs for much higher prices than lots of market segments are willing to pay.
I am also curious why would 4FMPAS be useful for Deep Learning when all the major performance improvements lately come from half 16-bit/quarter 8-bit-precision floating/fixed point math (you don't need precise boundaries).
This is a claim by Intel. See for example slide 140 of http://cs231n.stanford.edu/slides/2017/cs231n_2017_lecture15...
Smaller FP sizes mean you can run more of these in parallel. Smaller FP sizes mean (much) larger roundoff error, fewer bits of precision kept.
how is AVX-512 innovative when it's essentially SSE with 4x bigger registers? am i missing something?
It is much more: Just to give examples:
- The introduction of the opmask registers to mask many AVX-512 instructions
- AVX-512 Conflict Detection Instructions (CD) enables lots of loops to vectorize that could not vectorized before.
Also the fact that now 32 SIMD registered can be addressed does not seem very innovative from the outside, but implies a deep change: Before only 8 (32 bit mode; bite 5-3 of MOD-REG-R/M field) or 16 (64 bit mode; additionally use REX.R field) SIMD registers could be addresses for deep instructional encoding reasons. This also holds when using a VEX prefix. So being able to use 32 SIMD registers requires a completely new prefix scheme (EVEX). This new scheme contains lots of new capabilities (source: https://en.wikipedia.org/w/index.php?title=AVX-512&oldid=841...):
- Expanded register encoding allowing 32 512-bit registers.
- Support up to 4 operands.
- Adds 7 new opmask registers for masking most AVX-512 instructions.
- Adds a new scalar memory mode that automatically performs a broadcast.
- Adds room for explicit rounding control in each instruction.
- Adds a new compressed displacement memory addressing mode.
I'm really not convinced by this list.
But it's difficult to practically use AVX-512 when the turbo boost throttling prohibits breaking even with an equivalent AVX/AVX2 workload. I had a small project that would trivially scale up to AVX-512 register sizes. Despite doubling the vector width, it was actually much slower than the AVX/AVX2 version in practice -- just because it had such a low turbo boost ratio when running AVX-512 instructions.
This was also a problem with the first AVX implementations, so this is very typical: In the first generation, Intel makes such an instruction set extension available, so that one can write applications that make use of it (though they will usually not be faster, sometimes even slower). In the following generation this new feature is made fast, so that the newly written algorithms really get to profit from it.
Intel having significantly more vector capacity than AMD in each core is just a design decision, not an innovation.
Adding a single-precision FMA instruction is also nice but not particularly 'innovative'.
And it's not like ECC is innovative either. Intel is sitting on options, but it's not sitting on ideas.
No, AVX-512 is not only about doubling everything and calling it a day. I think the main innovation in AVX-512 is making it an easier vectorization target for compilers.
https://software.intel.com/en-us/articles/the-intel-advanced...
Intel's 10nm HVM isn't coming until 2019. That has been confirmed in an investor meeting when they were pushed to finally give an answer. So Intel has options? yes, lower Pricing, new SKU to muddle things up, more tricks like these Overclocking to damage AMD buyers expectation. Show unrealistic Cinebench that now totally flavours them. More Intel Inside tactics, changes of LLVM / GCC general code to force slower compile on AMD CPU, bundle their SSD with their CPU and force any vendor to get off AMD.
Oh yes, lots of options. None of them are technically superior though.
And btw, lots of the best Intel engineers are no longer with them. Apple, Qualcomm, Samsung, Nvidia, Telsa are all hiring their talents out of Intel Campus. If TSMC had Fabs in US I would bet many Intel Fabs engineer would have joined TSMC as well.
The problem with this kind of statement is that many people by their PC/laptop just for surfing and office applications. Especially for the latter application even a 7 year old PC currently suffices. But it is in my opinion easy to come up with ideas how applications that really need this kind of speed will profit from AVX-512. Additionally keep in mind that people who run applications that benefit from more speed will much more be prone to buy new PCs than people who just do office and surfing.
Infinity fabric seems quite novel for desktops.