1,083 karma · joined September 22, 2015
meet.hn/city/33.6856969,-117.8259810/Irvine
Or, put a dynamic factor into your vector size and design everything around it. Such that every platforms can plug in their own factor and _scale_ the size of vectors. This is basically what LLVM IR does for SVE and RVV: `<vscale x 4 x i32>` where vscale is the said dynamic factor. Though the exact value of vscale is only known during runtime, it doesn't matter -- we still can design compiler optimizations and lowering around it. The generated binaries can then be portable across platforms with different vscale values.
And the approach you mentioned here:
> a family of ISAs that are ABI compatible such that one can compile down to a semi-pre-optimized portable IR, and just do the last bit per ISA
I think this is basically WebAssembly and PTX, and one may argue, Java bytecode. Yet look at how much efforts and time it took for WASM runtimes and JVMs to actually produce performant machine code (the "last bit per ISA" you mentioned) for just a couple of architectures! (e.g. X86 and ARM). And I wouldn't surprised if NVIDIA pour even more money on building optimization pipeline from PTX to each of their different uArchs.
No
But it's really just an implementation difference, the idea is still to have a lightweight RTTI.
for one, its iterator won't be invalidated
Unlikely, as pointed out in sibling comments the core ISA is too limited. What might prevail is profiles, specifically profiles for application processors like RVA22U64 and RVA23U64, which the latter one makes a lot more sense IMHO.
> Even more of a wild idea is to pair up two cores and have them work together this way
I don't think that'll be profitable, because...
> When you have a core that would have been idle anyway
...you'll just schedule in another process. Modern OS rarely runs short on available tasks to run
A bit late to this comment but most of these intrinsics are overloads of different LMUL and SEW on a single instruction. I'm pretty sure the actual number of RVV instructions is way less. So maybe you could consolidate overloads of the same instruction into the same page or something.
yeah scheduling quality usually has a weaker connection to the performance of OoO cores. Though I would also like to point out:
1. in-order cores still heavily relies on scheduling quality
2. Issue width is actually a big thing in MachineScheduler regardless of in-order or out-of-order cores. So the problem you outlined above w.r.t different implementations of uops cracking is indeed quite relevant
3. MachineScheduler does not use the BufferSize -- which more or less mirrors the issue queue size of each pipe -- at all for out-of-order core. MicroOpBufferSize, which models the unified reservation station / ROB size, only got used in a really specific place. However, these parameters matter (much) more for llvm-mca> It doesn't support branches at all. I know it's a very hard problem, but that's the problem I have
Shameless self-plug: https://github.com/securesystemslab/LLVM-MCA-Daemon
[1]: https://android-review.googlesource.com/c/platform/packages/...
[2]: https://android-review.googlesource.com/c/platform/packages/...