Is there a possibility that in the not too distant future that GPUs and CPUs will just converge? Or are the tasks done by GPUs too specialized?
Is there a possibility that in the not too distant future that GPUs and CPUs will just converge? Or are the tasks done by GPUs too specialized?
CPUs aim to minimize latency (how many cycles have to pass before you can use a result), and do so by way of high clock frequencies, caches and fancy micro architectural tricks. This is what you want in most general computation cases where you don't have other work to do whilst you wait.
GPUs instead just context switch to a different thread whilst waiting on a result. They hide their latency by making parallelism as cheap as possible. You can have many more cores running at a lower clock frequency and be more efficient as a result. But this only works if you have enough parallelism to keep everything busy whilst waiting for things to finish on other threads. As it happens that's pretty common in large matrix computations done in machine learning, so they're pretty popular there.
Will they converge? I don't think so - they're fundamentally different design points. But it may well be that they get integrated at a much closer level than current designs, pushing the heterogeneous/dark silicon/accelerator direction to an extreme.
If you take a serial algorithm and put it on the GPU, it's easy to verify that a single GPU thread is much slower than a single thread on the CPU. For example, just do a bubble sort on the GPU with a single thread. I'm not even including the time to transfer data or read the result. You'll easily find the CPU is way faster.
The way you get GPU speed is by finding/designing algorithms that are massively parallel. There are lots of them. There are sorting solutions for example.
As an example, 100 cores * 32 execution units per core = 3200 / 20 = 160x faster than the CPU if you can figure out a parallel solution. But, not every problem can be solved with parallel solutions and if it can't then there's where the CPU wins.
It seems unlikely GPU threads will be as fast as CPU threads. They get their massive parallelism by being simpler.
That said, who knows what the future holds.
Here's mine
https://jsfiddle.net/jw7a6to9/ bubblesort
https://jsfiddle.net/y1w6s9tj/ taylor series
> in which case you can show those same GPU cores are 20x faster than CPU for other things
Which things? Remember, I wrote single thread, no SIMD, no samplers. It's the parallelism that provides the speed.
My dev GPU is a 6800XT. Cheapish gaming card from a little while ago, 16GB ram on the card. 72 "compute units" which are independent blocks of hardware containing memory ports, floating point unit, register file etc. Roughly "a core" from x64 world. Each of those can have up to 64 tasks ready to go, roughly a "hyperthread". It's 300W or so.
There's some noise in the details, e.g. the size of the register file from the perspective of a hyperthread affects how many can be resident on the compute unit ready to run, the memory hierarchy has extra layers in it. The vector unit is 256byte wide as opposed to 64byte wide on x64.
But if you wanted to run a web browser entirely on the GPU and were sufficiently bloody minded you'd get it done, with the CPU routing keyboard I/O to it and nothing else. If you want a process to sit on the GPU talking to the network and crunching numbers, don't need the x64 or arm host to do anything at all.