Meanwhile, the GPU cores will be simple cores with huge vector engines. This honestly seems somewhat similar to RDNA at a high level where you have a large SIMD unit with a scalar cores for branching and one-off calculations.
The big payoff is cohesiveness. Your RISC-V GPU shares the same memory model as your CPU. This should have a payoff in easy integration. Likewise, sharing a good permission model can probably help with all those GPU exploits in third-party code (like webGL).
Vector instructions have register size independent code which makes vector programming from CPUs much more approachable and binary compatible.
GPUs have geometrically relevant vectors as basic data types. I wish language designers would realize the value in that, but they are too concerned with abstraction.
RISC-V vectors are meant for the abstract general parallel concepts, but could also help with graphics code.
There really seems to be a disconnect around which types of vectors are good for which uses.
That is the naive application of SIMD, which many have attempted. It does not lead to a meaningful speedup.
To get a good speedup from SIMD, you need SOA or AOSOA layout. Ergonomically this is already similar to vector instructions.
How can it not provide a meaningful speedup? If I have 3- or 4-element vectors a,b, and c, with scalars u,v. I want to compute:
a = ub + vc;
These should be first class data types, passed by value, and that line of code should take 2 SIMD instructions at most. That should also not require a fancy vectorizing compiler to do the analysis to find the parallelism because the data types map directly to the ISA vector registers.
GPUs have this, the benefit should be very real.
The reason for this design is that while vec3 and vec4 are common, they are by no means universal in modern shader code. A micro-architecture built on vec3 or vec4 would be under-utilized in the large amounts of shader code that are scalar in the source language.
There are some exceptions, e.g. native support for 16-bit vec2 comes naturally on architectures with 32-bit registers. But those are exceptions.
It's the whole - full core on a GPU aspect, which probably isn't a bad idea for a few reasons, could even offload much of the driver onto it and with that, make platform drivers much more manageable...maybe.
This is an architectural level, not a microarchitectual level design (yet).