Anandtech: AMD Radeon HD 7970 Review (28nm, new architecture)
anandtech.com
anandtech.com
I'm finding it exceedingly hard to make any guess whether certain algorithms are worth trying to port over or not, because the explanations are almost incomprehensible.
To give an example, the VLIW->SIMD unit transition in this architecture compared to NVIDIA's scalar units.
As far as I understand the VLIW vs SIMD difference, a VLIW instruction is more powerful compared to SIMD because one "long instruction" can contain different operations on different data. Whereas SIMD is the same operation on multiple data.
Traditionally, VLIW was entirely statically scheduled, which put all the burden on the compiler. Because graphics cards recompile all shaders anyway, it's not a bad fit.
So, now AMD changed their units from 16x4 VLIW to groups of 16xSIMD engines. The advantage here is that you no longer have to have groups of 64 similar ALU operations, but can do with groups of 16 similar ALU operations. Conversely, there should be more of such groups, i.e. more control logic compared to the old design. To top that off, there's improvements that allow one to schedule multiple threads over the GPU at once.
Am I on the right track here?
If I am, whats the minimum amount of identical operations that you must do to achieve reasonable throughput? 16 identical operations for 100% efficiency? Put differently, if I have code that requires different operations (due to branches, i.e. effectively conditional operations) on each computation stream that I'm pulling through, what factors are going to limit the effective speed I get out of this?
VLIW exists to exploit very wide instruction parallelism while clearly delineating what instructions have dependencies on other instructions, and (at least on the Radeon) usually end on a memory write or a complex branch.
Radeon clauses are up to 128 instructions long, and manage 5 ALUs (on 4xxx, 5xxx, and 68xx, 4 of them identical, 1 of them being able to do double precision math and transcendentals), or 4 ALUs (on 69xx, all 4 identical and does not require a specialized ALU for DP and T).
The compiler optimizes dependency flow and ALU usage, instead of requiring dedicated hardware common in CPU design. This means far less silicon is dedicated to the task, and instruction scheduling is far more predictable and optimized.
Your suggestion that AMD used 16 ALUs per pipe under VLIW is wrong, The change is 4/5 VLIW ALUs to 16 ALUs that now can execute SIMD instructions. The compute units (the head end that synchronizes multiple pipes to perform one task in parallel) still use VLIW-like clauses to synchronize pipe usage.
Your suggestion that this new arch allows you to schedule multiple threads on the GPU at once is nonsensical: the correct term for pipe is "hardware thread", on a GPU like 5870 you have 320 hardware threads (1600 ALUs), you already schedule all of them at the same time for massively parallel execution.
What has changed is the CUs now support running clauses from different shaders at the same time by using some CU on one task, and some on another, and I believe it may also be able to have clauses from different shaders loaded at the same time and switch without overhead; on VLIW Radeons, shader change out has a high context switch penalty.
The only thing the new GCN arch really does is allow the ALUs to operate on SIMD instructions which allows higher instruction packing. This does not mean they do not use VLIW-like clauses, and it doesn't mean it is like Nvidia's design (which the media keeps repeating).
Nvidia's ALUs are free form stream processors, and do not have a clear beginning or end to each clause (as the hardware does not exploit hardware thread synchronization), they frequently suffer from cache misses and pipeline stalls, and they cannot easily exploit instruction level parallelism.
In addition, Nvidia does not exploit deep pipelining. On VLIW Radeons, the instruction pipeline is multistaged and 4 instructions deep, so by the time you are submitting the 5th instruction you are getting the results from the first and instructions 2, 3, and 4 are still being processed. This allows much easier synchronization between ALUs since they all read/write to the same set of registers.
The addition of SIMD instructions to this design allows much higher data throughput and much higher instruction packing; instead of executing, say, two Bitcoin hashes per VLIW4/5 group, each instruction winding around the group for maximum ALU efficiency, you can run 4 (or however many GCN uses for SIMD, most likely 4 or 8) hashes at the same time as a SIMD operation and not require complex compiler maneuvering to do a clearly instruction parallel operation (thus 16x4 hashes per group).
Now, ultimately, nothing of what I've written actually matters. OpenCL is a black box on purpose, it doesn't matter how the implementation executes it as long as it does so correctly and efficiently. AMD is betting that GCN is more efficient for the given silicon real estate.
Actually, the previous architecture was a setup of 16x simd, where each of the simd operations was a 5/4 wide vliw. So calling that 320 hardware threads is wrong -- in Cypress there really was only 20 front-ends which drove these bundles of 80 alus in groups of 5x16. Also, it was a 4-long barrel processor, so you had to schedule a SIMD "wavefront" of 64 "threads" for each unit.
In the new version each CU still has 4x16 ALUS like it had in Cayman, but now each of the 4 simd units of 16 elements can be scheduled independently by a different hardware thread.
The R700 Programming Manual seems to indicate my interpretation is correct, although if you can provide evidence that I'm misinterpreting it, I'm all ears.
The barrel is essentially used to extend the vector registers from 16-elem to 64-elem, and a 64 "thread" wavefront, consisting of 5 VLIW'd instructions is essentially the smallest amount of work that R700 can do.
I only "speak" CUDA, however the primitives provided by the CUDA/OpenCL development platforms are largely equivalent. For a single "kernel" the GPU provides parallelism at two levels,
1. Fine-grained Thread-level parallelism: You have a group of threads performing ideally identical operations. The model is called SIMT (Single Instruction Multiple Thread) because it is not quite SIMD but close. Say you have (1-1024) threads, the hardware splits these into warps of 32 threads. The 32 threads within a warp execute together in classic SIMD manner, however only one of these warps is executing on the actual hardware at any given time. When a warp makes a request for data from high-latency device memory, it is typically suspended and another warp swapped in. This allows latency hiding between warps. However, calling __syncthreads() forces all warps to synchronise. Any conditional statements cause warp divergence, the warp is split and the alternative execution paths run sequentially. Clearly, this incurs a performance penalty as SIMD behaviour is lost and should be avoided as much as possible. The 'ideal' identical operation size is therefore a warp or 32 consecutive threads, however occasional divergence is acceptable. Memory access patterns are much more important however.
2. Coarse-grained Block-level parallelism. A block is an independent group of threads. The platform allows mutiple blocks to execute on different Streaming Multiprocessors. The hardware may schedule multiple blocks to execute on a single SM if there are sufficient resources (such as register and shared memory), this also facilitates latency hiding. Other than the kernel invocation which created multiple threads blocks, they are unrelated and cannot communicate or be synchronised.
The potential thread level parallelism is what should be considered. If your algorithm is representable as fine-grained floating point vector operations without excessive divergence then it is well-suited to the GPU. Of course, the actual performance gain depends on the quality of implementation and optimisation is highly GPU-specific and non-trivial. I recommend
http://developer.download.nvidia.com/compute/DevZone/C/html/...
as an example of how important optimisations lead to those impressive performance numbers.
For AMD’s technology each tile will be 64KB, which for an uncompressed 32bit texture would be enough room for a 4K x 4K chunk.
Isn't this off by a factor of 1,000? 4K x 4K is 16M texels, which at 32 bits per texel would require 64 MB. A chunk of 64 KB cannot hold that. They repeat the "64KB" value for the chunk size many times, not sure in what direction they're wrong, really. I guess if I kept up more with graphics tech, the answer would be obvious. :)
64k can fit 128x128 rgba texture, which looks like the correct size if you look at the "PRT Translation table" image.
A: Not all that many which is why breaking that up into smaller chunks an enabling a virtual 4k texture without compression is a good idea.
Do not forget that mining bitcoins becomes progressively harder with time.
Thus to get any sort of meaningful reviews, they would have to review whole field of graphic cards each time a new one comes out.
Not feasible.
Edit: Ok, so I don't know anything about Bitcoin mining, thanks for clarifications.
Knowing how fast a card is at bitcoin mining is flatly uninteresting to anyone who is not using it as one.
Also, the data processed is around 104 bytes (two SHA265 executions: the second part of the first execution for the 80 byte header (hash state from first part + the next 16 bytes) (the first part of the first execution only needs to be done once and is done on the CPU), and then the second execution (hashing the result of the first execution)).
Even if you're in stereo (1/120) and you write a new image to the buffer 100 times per frame (1/12000) that still gives you 40,000 cycles on a 500MHz clock.
No really, when we will get something as simple as OpenMP to do our data mining on GPUs? There is more data to be processed than there are games to play.
Not that that's really necessary, I mean the complexity of scientific computing (parallelism, numerical stability, etc.) means the programmers that do it have no trouble with CUDA et al (which is fairly simple if you just forget about coherency).
OpenMP is considering adding accelerator support in 4.0. In the meantime, NVIDIA and friends are having a go at specifying semi-standard pragma-based accelerator support with OpenACC.