RV64X: A Free, open-source GPU for RISC-V
eetimes.com
eetimes.com
The name is slightly reminiscent of https://en.wikipedia.org/wiki/File:RIVA_TNT2_VANTA_GPU.jpg
There's been one major success case: A64Fx and ARM SVE. The Fujitsu processor just recently took #1 in LINPACK, #1 Green500 (also the most efficient) and also #1 in HPCG (an alternative benchmark that GPUs traditionally suck at).
Its not a full GPU yet, but that A64Fx is clearly a powerful SIMD compute cluster, very similar to a GPU. It has succeeded where the Xeon Phi seems to have failed.
That said, if the goal is to compete with existing GPUs for general compute workload, the graphics part is unnecessary anyway. That could work out, though calling it a GPU is a red herring.
Blending pixels is becoming less and less important in the power budget of mobile devices too.
Again, this doesn't matter for a pure compute part. It's just not particularly clear from the article whether they're even aware of the distinction.
I guess it makes sense; the RISC-V crowd might possibly skew towards early adoption over backwards-compatibility.
It would be really cool to see some implementations from places like SiFive/GigaDevice/etc. Imagine how easy driver support could be if everyone used and contributed to the same open IPs...
I expect that, in the future, gl and d3d will be implemented entirely on top of vulkan, and thus be more portable and make it easier to make GPUs and graphics drivers.
I'm not sure whether you're talking about the "official" Microsoft D3D libs here, but I very much doubt that'll ever happen. They don't like dependencies they can't control. The Excel team used to maintain their own compiler because they didn't want to be dependent on the Visual C++ team who worked in the same building.
Same with Apple and Metal.
Vulkan is not even supported on UWP or Win32 sandboxes, the ICD mechanism its drivers use from the OpenGL days is only allowed in classical Win32 mode.
Version 3.0 is the latest on iDevices and while Android can do up to 3.2, it is an optional API.
Likewise, Metal is the name of the game on iOS nowadays, and Vulkan was introduced in Android 7 as optional API, and only became mandatory in Android 10.
Vulkan also carries on the tradition of Khronos APIs, extension spaghetti, so while a device might support Vulkan, it doesn't mean it supports the Vulkan that the application actually needs.
It's the whole - full core on a GPU aspect, which probably isn't a bad idea for a few reasons, could even offload much of the driver onto it and with that, make platform drivers much more manageable...maybe.
This is an architectural level, not a microarchitectual level design (yet).
Vector instructions have register size independent code which makes vector programming from CPUs much more approachable and binary compatible.
GPUs have geometrically relevant vectors as basic data types. I wish language designers would realize the value in that, but they are too concerned with abstraction.
RISC-V vectors are meant for the abstract general parallel concepts, but could also help with graphics code.
There really seems to be a disconnect around which types of vectors are good for which uses.
That is the naive application of SIMD, which many have attempted. It does not lead to a meaningful speedup.
To get a good speedup from SIMD, you need SOA or AOSOA layout. Ergonomically this is already similar to vector instructions.
How can it not provide a meaningful speedup? If I have 3- or 4-element vectors a,b, and c, with scalars u,v. I want to compute:
a = ub + vc;
These should be first class data types, passed by value, and that line of code should take 2 SIMD instructions at most. That should also not require a fancy vectorizing compiler to do the analysis to find the parallelism because the data types map directly to the ISA vector registers.
GPUs have this, the benefit should be very real.
The reason for this design is that while vec3 and vec4 are common, they are by no means universal in modern shader code. A micro-architecture built on vec3 or vec4 would be under-utilized in the large amounts of shader code that are scalar in the source language.
There are some exceptions, e.g. native support for 16-bit vec2 comes naturally on architectures with 32-bit registers. But those are exceptions.
Meanwhile, the GPU cores will be simple cores with huge vector engines. This honestly seems somewhat similar to RDNA at a high level where you have a large SIMD unit with a scalar cores for branching and one-off calculations.
The big payoff is cohesiveness. Your RISC-V GPU shares the same memory model as your CPU. This should have a payoff in easy integration. Likewise, sharing a good permission model can probably help with all those GPU exploits in third-party code (like webGL).
Anyways I expect RISC-V to appear in Chinese mobile phones before laptops as the main unit.
It would be fun to run Android on this RV64X architecture....I guess that's the goal.
Cheap phones running Android would be a great start for RV64/RV64X though.
- Focusing on one use case might make it difficult to generalize the design later. Lessons learned might not carry over, and the platform will get anchored into particular design choices. This only gets worse the more the solution gets deployed.
- Having a product that can do both makes it simpler to do applications that combine multiple computational loads. Else, the data has to be transferred to another device. This seems unattractive since graphics, ML and massively parallel scientific computations are more similar to each other than to what CPUs are used for.
- Development budget gets fragmented, which is a disadvantage since the different use cases might not correspond to markets that are large enough to support each product.
- One of the platforms might just die off. Users that use both would be stuck in between and would have to find a solution, and most likely it will be NVidia/AMD again.
https://upload.wikimedia.org/wikipedia/commons/f/f8/DIAMONDS...
https://upload.wikimedia.org/wikipedia/commons/6/6b/Matrox_M...
http://vgamuseum.info/images/palcal/profi/665_elsa_xl_top_hq...
I hope this will not happen, but especially certain laptop companies always wanted to have non upgradeable RAM and only stepped back because users where to unhappy about it. But now they will point to Apple and tell you that's necessary for fast low power laptops with long battery live.
Do you think we could extend the memory of most 8 and 16 bit home computers?
If we wanted to upgrade our computers, beyond plugging stuff on the external expansion port, we had to buy a new model.
The PC was the precedence, and only due to IBM losing control of it.
Now with laptops and tablets becoming the standard consumer computers, we are back into the 80-90's computer form models, just instead of plugging them into the TV, the screen comes along for the ride.
And actually it was great, because it meant developers had to learn to extract all the juice of the computers people had, instead of expecting us to spend money upgrading.
Do you think we could extend the memory of most 8 and 16 bit home computers?
I'm not sure how your defining "most" home computers, but the c64/128 & apple ][ line were some of the most popular and they definitely had ram upgrades.https://www.c64-wiki.com/wiki/Commodore_REU
The soldered laptop ram thing, is fairly recent and by no means universal. I suspect it would be rarer than it is, if retailers put a little note on the machine descriptions "upgradeable RAM/upgradable disk" but they can't even be bothered to put the cpu clock rates (or sometimes even the core count) on the machine descriptions.
Which I do confess never to have seen, as C64 and Apple weren't anywhere to be seen in the Iberian Peninsula.
OTOH, the trend has been to put HBM/edram/etc on package for a while. Various intel/etc products have done that, which provides a fast tier of memory that you put in a local NUMA domain. Then put the actual DRAM in its own domain.
The problem so far is that strategy doesn't actually work well for most applications because they aren't sufficiently NUMA aware to take advantage of it. Instead what ends up happening is the "fast" ram ends up being a LLC tier.
It all ends up sorta being a physics thing though, it seems as you increase capacity, distance tends to grow as well. Meaning that there will be a limit to how much RAM they can toss on package and still gain a latency advantage over just soldering it to the board.