The MRISC32 – A vector first CPU design
bitsnbites.eu
bitsnbites.eu
https://pdfs.semanticscholar.org/b9e8/fcf11b662a31cecd5a08d5... https://www.sigarch.org/simd-instructions-considered-harmful...
In the simplest possible design, the cost of vector support in the MRISC32 is essentially the cost of the register file (i.e. the register memory), since the vector control logic just consists of a few adders and flip-flops, and the scalar execution units can be reused for vector operations.
You can also play other games like internal recoding the FP format if you split the register files.
On the other hand, you can remove the need for IntToFP and FPToInt move instructions if you use one register file.
At the end of the day though, I suspect it largely comes down to having more addressable architectural state to keep the two register files separate. Why have only 32 scalar registers when you can have 32 scalar int and 32 scalar fp registers?
On the MRISC32 you have a vector register file with 32 registers that can easily be configured to do scalar operations (integer or floating point) by setting the vector length to 1.
Have you looked at the Convex C-Series architecture? They copied the Cray vector idea. Starting with the C2 (I think) there was also a Vector Mask (VM) register which described which vector elements had valid data. Since the C-Series (and the Cray, I think) processed vectors serially, vector ops using the VM would only spend time on the valid elements. It was targeted at codes like this:
for (i = 0; i < n; i++) {
if (A[i] < 34) {
D[i] = A[i] * B[i] + C[i];
}
}
That would translate into code like this. (Actual instructions and mnemonics lost to the mists of time; this is pseudo-assembler.) ld vl, #N ; assume N <= 128 (-:
ld.v v0, A ; load A
ld s0, #32
less.v v0, s0 ; stores boolean vector into VM
ld.v.t v1, B ; load B where mask == true
mul.v.t v3, v1, v0 ; calc A*B under mask, store in v3
ld.v.t v2, C ; load C under mask
add.v.t v3, v3, v2 ; calc A*B+C under mask, store in v3
st.v.t v3, D
I am ignoring the difference between I32, I64, F32 and F64 because I don't remember how those were coded into the mnemonics, sorry.There were also instructions to load and store the VM to a scalar register pair.
The Mill architecture, if I understood the lectures, only has a masked store instruction; other vector instructions calculate the values for all elements and maintain an "invalid result" bitvector. An exception is only triggered if an invalid result is actually stored.
While the MRISC32-A1 does serial vector processing, the idea is that you should be able to do parallel processing with the same ISA, so I'm not certain that masking in that fashion is a good option for the MRISC32. Also, having a vector mask register like the Cray limits the size of the vector registers (e.g. to 32 elements in the MRISC32, or 64 elements in a 64-bit architecture).
I don't understand what you mean by this? The Mill team seem intent of producing a chip of their own.
https://github.com/mbitsnbites/mrisc32/tree/master/mrisc32-a...