Yet due to sheer clock speed increases and HBM improvements (16GB, 1 TB/s, wow) they actually seem surpisingly competitive. I'm both incredibly impressed and incredibly underwhelmed.
Yet due to sheer clock speed increases and HBM improvements (16GB, 1 TB/s, wow) they actually seem surpisingly competitive. I'm both incredibly impressed and incredibly underwhelmed.
60 CUs at 1.8GHz is way faster than AMD Fury ever was. Not only do you get the GHz advantage, but more L2 cache, more HBM2 Bandwidth, and architectural advancements (various instruction set improvements like double-speed FP16 support)
EDIT: AMD Fury's main advantage was that 1GHz clock made all of the nanosecond-level timings super easy to understand. 10-clock cycles was 10-nano seconds. Hurrah!
16gb seems like overkill for the next couple years of game releases.
It’s essentially just to keep up some production volume this is their HPC die sold to the consumer market its the Titan V equivalent die of AMD minus the Tensor Cores just for much much cheaper.
It’s likely going to be a flop gaming wise, 2x8 pin power connectors put it at 300W TDP with performance only matching that of a 1089ti/2080 in the best case scenarios AMD currently has.
The 1080ti can be had for under $600 if you can find it the 2080 is around $700 but you get RTX and DLSS which while isn’t that useful yet is going to be more useful for gaming than 16GB of memory.
GCN’s vector machine doesn’t scale as well as NVIDIA’s scalar architecture NVIDIA can add cores without caring about concurrency or ILP, since each core is an individually addressable scalar ALU, this is why “async” compute doesn’t benefit them as much as it does GCN cards that have a 4 wide SIMD array that is nearly as hard to feed as their old VILW arch and a massive external compute scheduler that sits idle when doing graphics.
> NVIDIA’s scalar architecture
NVidia doesn't have a scalar architecture in any meaningful sense of the word. Individual work items execute in warps, meaning (up to) 32 items execute the same instruction, at the same PC. If you have divergence, i.e. fewer than 32 items at the same instruction, you lose ALU throughput. This is the same as in AMD (and every other GPU architecture out there, for that matter).
The difference to AMD's architecture since Volta is that there is hardware support for each SIMD lane having its own PC, which makes it easier to have independent forward progress and a bunch of other features. The separate lanes don't have separate instruction fetch, though -- they all execute the same instruction (and if they can't, then some lanes will be disabled).
> this is why “async” compute doesn’t benefit them as much as it does GCN cards
Nvidia literally cannot do async compute in the same way as AMD, because their micro-architecture cannot switch between graphics and compute on as finely grained a basis as GCN cards. That's why they don't benefit from async compute. (Though I'm not sure whether that's still true since Volta.)
> GCN cards that have a 4 wide SIMD array
GCN has 64-wide waves, though with a bit of a weird execution scheduling which means that 16 lanes out of the 64 are computing at a time. Nowhere in GCN are there any 4-wide SIMD arrays.
> a massive external compute scheduler that sits idle when doing graphics.
Also nonsense. Both AMD and Nvidia have the majority of die space allocated to the compute cores (whatever they're calling them). The pieces of fixed function logic that distribute work onto the cores (whether for graphics or compute) are comparatively tiny.
CUDA threads are scalar, each CUDA core is a single scalar ALU, which CUDA cores are assigned to a warp is flexible.
>Nvidia literally cannot do async compute in the same way as AMD, because their micro-architecture cannot switch between graphics and compute on as finely grained a basis as GCN cards.
Both AMD and NVIDIA GPUs have to do context switching. The difference is in scheduling GCN has a dedicate scheduler with 8 queues only for compute that is the ACE it sits utterly idle while doing graphics the graphics scheduler sits within each CU.
>Nowhere in GCN are there any 4-wide SIMD arrays.
Each CU contains an array of 4 SIMD units which each take a vec4 input, GCN is much closer to VILW4 than you think. Each CU also contains a scalar unit it cannot excute them in parallel.
If the SIMD units cannot be fed they sit idle, idle CUDA cores can be assigned to a different warp.
>Also nonsense. Both AMD and Nvidia have the majority of die space allocated to the compute cores
ACE takes up about 15% of the die space that is big in my book.
> Each CU contains an array of 4 SIMD units which each take a vec4 input
AMD SIMDs are (poorly named) scalar units, which are programmed using scalar arithmetic. FP16 is the only exception, as FP16 is a true SIMD operation (taking place inside a AMD scalar SIMD 2-at-a-time).
Yeah, AMD really needed to name their stuff better. "vGPRs" ("vector General Purpose Registers") are in fact scalar units, as documented in the Vega Instruction Set.
See chapter 6 for proof: https://developer.amd.com/wp-content/resources/Vega_Shader_I...
The SIMDs replaced historical SIMD-units from the 6000-series roughly 10 years ago. I bet that the name is historical in nature. But AMD's cards are scalarly programmed, just like NVidia's. There is no need to use vec4 on modern (aka: anything in the last 10 years) cards from AMD. Half2 is the only thing that gets a benefit.
> Each CU also contains a scalar unit it cannot excute them in parallel.
The "Scalar" unit in AMD GPUs is... closer to a boolean-vector unit. AMD really chose bad names for these things... But anyway, it is a 64-bit unit, where each bit is used to calculate the execution masks to the vector units.
The "Scalar" unit runs once every clock-tick. "Vector" units repeat themselves over 4-clock ticks (and therefore execute 4x slower than the Scalar unit). Ultimately, 1-Scalar unit can perform roughly the same amount of work as the 4-vector units.
See page 31 for a great example on how the sALU of Vega is used to create loops and handle divergent cases. The "sALU" is effectively the unit to calculate "if" statements and "loops".
This is a really cool way to put it, though I'd point out that the scalar unit is actually also used for (uniform) arithmetic, especially on pointers / indices. For example, if you have memory accesses of the form uniform_base + stride * thread_id, you can use the scalar ALU to compute / manipulate the uniform_base part.
If you think of the GCN compute units as CPU cores with a really wide SIMD + powerful scatter/gather + texture sampling unit, the chosen terminology of scalar/vector ALUs makes sense.
The difficulty is in mapping the (work item-centric) programming model that developers see to the hardware, but that's conceptually fairly similar to what ISPC does.
I have a feeling you're confusing lanes vs. cores here. Individual "threads" (work items) cannot migrate between warps, and every "core" only issues one instruction at a time (per execution port, in some of Nvidia's architectures). Each execution port is a SIMD execution port executing that instruction for multiple work items. If the work items of a warp diverge, some of the lanes of those SIMD execution units will be idle.
This is the exact same behavior as on AMD's GCN, though obviously the terminology is different and some of the other details may be wired differently.
> ACE takes up about 15% of the die space that is big in my book.
I have a feeling that you're just making up numbers here. Can you point to evidence?
While I agree with your post in general, I'm not totally in agreement here. If you look at the context in which G80 and CUDA were introduced, it has a significant scalar dimension: for a single "thread" writing an addition between two float3 ended up being executed as 3 separate additions on the scalar components of the vector. This was, if I remember correctly, a departure from previous architecture optimized for graphic processing. In this sense the ISA is/was mostly scalar.
If you look at the actual hardware, though, they're clearly both vector architectures: each "core" has N SIMD-style lanes executing the same instruction simultaneously; it's just that each of those lanes corresponds to a different work item of the compute dispatch.
This applies equally to AMD's and Nvidia's architectures.