I'm sure Intel should fix the problems Linus is complaining about, but I feel like chip vendors are being forced into this "add special purpose blocks" approach, as the only way to make their new chips better than their old ones.
I'm sure Intel should fix the problems Linus is complaining about, but I feel like chip vendors are being forced into this "add special purpose blocks" approach, as the only way to make their new chips better than their old ones.
Dealing with these issues might require you to know the corners of the instruction set really well or some times the solution is outside of the instruction set and is related to how your data structure is laid out in memory leading you to AoS vs SoA analysis etc.
Compilers and vectorization: Based on reading a lot of assembly output I think what compilers usually struggle with are assumptions that the human programmer know hold for a given piece of code, but the compiler has no right to make. Some of this is basic alignment, gcc and clang have intrinsics for these. Some times it's related to the memory model of the programming language disallowing a load or a store at specific points.
GPGPU programmability: GPUs being easy to program is something I take with a grain of salt, yes it's easy to get up and running with CUDA. Making an _efficient_ CUDA program however is easily as challenging if not more than writing an efficient AVX program.
https://pharr.org/matt/blog/2018/04/18/ispc-origins.html#aut...
> as long as vectorization can fail (and it will), […] you must come to deeply understand the auto-vectorizer. […] This is a horrible way to program; it’s all alchemy and guesswork and you need to become deeply specialized about the nuances of a single compiler’s implementation
It's not all that different conceptually to AVX-512 with mask registers, except the vector size is even larger and of course the programming model differs.
I have a simplistic explanation - maybe not what you're looking for but it is the best I can do...
At 12m23s in the video he says, "If you're working in a layer and the layers are well constructed (abstracted) you really can make a lot of progress. But if the top layer says, 'to make this really fast, go change the bottom layer', then its going to get all tangled up."
That's what implementing an algorithm on a SIMD architecture feels like to me. I have to figure out a way of filling my SIMD width with data each clock cycle, while in contrast, the specification of the algorithm deals with data one piece at a time.
Take insertion sort as a (bad) example.
i ? 1
while i < length(A)
j ? i
while j > 0 and A[j-1] > A[j]
swap A[j] and A[j-1]
j ? j - 1
end while
i ? i + 1
end while
That algorithm cannot easily take advantage of SIMD. You have to change the algorithm to make it work with the architecture.We'd probably say the algorithm is the top level of the abstraction stack, and the SIMD architecture is a level near the bottom. So this problem is the opposite way around to how Jim phrased it, but the point is that we have NOT got clean abstraction - an implementation in one layer depends on the implementation in another.
I prefer intrinsics as they give more control than shader languages and they can be written in C++ instead of fiddling with some garbage GPU API that runs async.
Also one of the reasons CUDA won developer love is that it fully embraced polyglot programming on the GPU.
CUDA seems nice, but being Nvidia only makes it a total dead end.
There's also HIP[1], which can be used as a thin wrapper around CUDA, or with the ROCm backend on AMD platforms. It doesn't yet match CUDA in either breadth of features or maturity, but it's getting closer every day.
I wish all the GPU companies would get together and make a standard based on C++ and stick with it.
In what concerns commercial uses of CUDA, Hollywood doesn't seem to have any problem with it, nor the car manufacturers with Jetson.
It does look like Intel is supporting at least, so maybe in the future it will be a good option.
Or are you speaking about the 1% Linux users on Steam?
Kids these days get 8 cores for a 100W TDP.
When I was a boy, 100W got you a single core. And you didn't get dynamic frequency scaling, so it'd be putting out that heat all the time.
(We also had to walk to school barefoot in the snow, uphill both ways)
386, introduced 1985:
http://www.cpu-world.com/CPUs/80386/Intel-A80386-16.html
Typical/Maximum power dissipation: 1.85 Watt / 2.3 Watt
And even no Pentium III 1999-2003 needed more than around 30 W:
https://en.wikipedia.org/wiki/List_of_Intel_Pentium_III_micr...
Even if the frequency was fixed, dissipated heat did definitely vary together with the computing load.
What's the problem? My old school pentiums kept my dorm room nice and toasty. Could keep my window cracked in the winter for fresh air while gentoo compiled...
Because outside artificial intelligence, graphics and audio, there is little else that common applications would use the GPGPU for, so the large majority of software developers keeps ignoring heterogeneous programming models.
With how AVX512 is implemented, there isn't much point in a compiler auto optimizing general purpose code to use it, because even if there is a theoretical speedup, it may well be slower in practice.
No. I recently could really, really have used the packed saturated integer arithmetic and horizontal addition in AVX2 (but my old machine doesn't support it) and even better, the same but 512 bits wide on AVX512. It would only have been 6 or 7 instructions, if that, but it was inner loop, and mattered. Using compiler intrinsics would have been fine. I think you're looking at things too narrowly.
Actually we already have openmp to cuda (http://www2.engr.arizona.edu/~ece569a/Readings/GPU_Papers/3....) so just making it more production-ready would be perfect.
I think you got this backwards - the lack of developers' interest is what leads to the mistaken impression that GPU compute is only good for multimedia and FP-crunching workloads. Even looking at the success of GPU compute in mining cryptocoins (only ASIC's do better) ought to be enough to tell you that we could do a lot more with them if we cared to.
That is simply not true. You can run the 64 Core on EPYC 2 all at once at 3Ghz all with Air Cooling.
At every node they have reduced power consumption that is also one reason you see continuous performance improvement.
I'm not claiming anything controversial. Power not having scaled as well as area recently is often referred to as the end of Dennard scaling:
https://en.wikipedia.org/wiki/Dennard_scaling#Breakdown_of_D...
> You can run the 64 Core on EPYC 2 all at once at 3Ghz all with Air Cooling.
That can be true despite the fact that power hasn't scaled as well as area.
> At every node they have reduced power consumption
Yep, just not as much as they improved area.
The "weak form" of Moore's Law--"Performance doubles every 12-18 months"--is dead and buried.
The "strong form" of Moore's Law is still active--"Transistor cost halves every 12-18 months".
This means that you can't make the primary paths any faster. So, all you can do is add functionality and pray that someone magically can make that functionality relevant to the primary use cases.
Crypto or video decoding comes to mind, those would be much faster with dedicated silicon, but more general AVX instructions can get you halfway there. Well, maybe a quarter. People point out that AVX uses a lot of power, but they ignore that the same algorithm running instead on more but simpler cores would use even more power.
Maybe misunderstand you but there are some fairly non-general ops for encoding/decoding crypto
That is the point of Linus. He would have preferred to use that increase in transistor count for other things, like more cache.
Skylake is less than 30% cache. However internally it's 512bus, thanks to avx-512 - which could be considered suboptimal.