SiFive Rolls Out RISC-V Cores Aimed at Generative AI and ML
allaboutcircuits.com
allaboutcircuits.com
https://www.semianalysis.com/p/sifive-powers-google-tpu-nasa...
What are push/pop vector instructions? Is that kind of like the old school OpenGL 1 era matrix stacks? Or literally just traditional stack style access between memory and a vector register file? Or something else entirely?
These are CPUs with Vector 1.0. Details will follow in the actual announcement, tomorrow.
0. https://twitter.com/sophgotech/status/1713852202970935334
Just port/extend pytorch to make it easy to utilize fused ops.
IMO - you’d have better luck as a hardware vendor implementing an LLM toolchain and bypassing a general purpose DL framework. At the very least you should be able to post impressive results with this approach rather than a half baked pytorch port.
Bigger problem for startups trying to muscle in on LLMs is that there isn't much room for improvement on existing solutions to do something radically different.
aye - unless you are able to notch a 10x cost/performance improvement. The migration overhead will just make it not worth it to switch.
Say you took all the effort in the world to build your custom LLM toolchain to train a Llama on custom hardware. And then suddenly someone comes up with LoRA. You didn't even finish porting it to your toolkit then someone comes up with GPTQ.
Can't keep up with a custom toolchain imo.
It's like a forked linux kernel. Eventually you're gonna have to upstream if you're serious about it, which is what AMD is actively doing with pytorch for ROCm (masquerading it as CUDA for compatibility).
To be fair, the M1/M2 chip can't be purchased or used separately from the Mac, unlike GPUs or socketed CPUs, and demand for Macs is already fairly high.
https://www.youtube.com/@GroqInc/videos
Other interesting Groq tidbits - their models are deterministic, the whole system up to thousands of chips runs in sync on the same clock, memory access and network are directly controlled without any caches or intermediaries so they also run deterministically.
That speeds up communication and allows automatic synchronisation across thousands of chips running as one single large chip. The compiler does all the orchestration/optimisation. They can predict the exact performance of an architecture from compile time.
What makes Groq different is that they started from the compiler, and only later designed the hardware.
All the big chip startups have their own pytorch compiler that works on the examples they write themselves. From what I've seen of Groq it doesn't appear to be any different.
The problem is that pytorch is incredibly permissive in what it lets users do. torch.compile is itself very new and far from optimal.
But there's a real technological change going on. Machines are capable of performing tasks that were nearly unthinkable two years ago. There's real progress beneath the hype.
AI workloads are going to be dramatically more significant in the next decade. A shift like this opens up real opportunities for smaller players to capture enormous market share. If SiFive manages to develop the best AI hardware, they could become bigger than Apple. Or they could fumble and go out of business. Either of these outcomes is genuinely possible when this kind of technological shift happens.
Like what? What's unthinkable tasks are so useful thanks to AI?
Speech -> immersive vr is an engineering challenge at this point.
They may be worth more than Arm, Ltd. in a decade or so. That also depends upon Arm making some more bone-headed decisions too.
Erasing strangers or unwanted items in the background on pictures taken by my phone locally from my Phone's photo app was unthinkable a few years since you needed Photoshop on a PC for that. I assume the TPU is doing heavy lifting on image segmentation since the UX is pretty much touch-to-erase.
https://i.imgur.com/y5BfxBF.jpg
Two years ago this immensely useful feat would have been impossible to accomplish in 30 seconds.
On a more serious note, ChatGPT has made made me a much more productive person. Recent example where it helped me a lot in getting an understanding of a new topic: https://news.ycombinator.com/item?id=37622335 - Even when it's making a fair amount of mistakes, it can be like a personal tutor who points you into the right directions when you're stumbling around in an unfamiliar field. Sort of like a much smarter search engine.
Granted, useful to me doesn't mean useful to you. But there's also people to whom search engines aren't useful, because they don't know how or when to use them.
Theranos is a great example. Good number of rich jumped into the tech while those in the industry knew a drop of blood does not contain the proper sample size for testing multiple markers.
Dr. Varun Sivaram's book "Taming the Sun" even talks about Elon Musk's solar panel grifting tactics to seek and exploit investors.
How is this any different than investors going for the world changing technology such as blockchains and NFTs?
*Edit, fixed spelling.
For extension that are implemented, better if there's a standard.
As for the # of them: just look what's been bolted onto x86 over the years, and their counterparts in other ISAs like Arm.
Consolidation will happen in the market, over time. Some combinations of extensions will be common, some rare. And due to the modularity, not a big issue for software.
The thing is, for some things you can get 100x improvement. Even if not used often its gone make a difference.
But for them most part, things that are part of RISC-V RVA23 profile make sense to be on a standard PC or Server.
Asked differently what isn't really necessary that is in RISC-V? One could make arguments for a simpler vector extensions for example. And a few other tiny things you can argue about.
Overall RISC-V RVA23 is a reasonable and feature complete set for a modern processor.
x86 on the other-hand has quite a few things that are not really needed if you redesigned it.
At least for stuff like vector extensions and tensor units, which could easily be emulated in software, this would seem to be a desirable trait. And I mean, it is hard to believe that most of this stuff isn’t proof-of-concept tested using software implementations first.
In particular, it could be nice for people who want to develop on RISC-V, when really good desktop chips come out, to be able to actually compile their whole project and run it before sending it off to the exotic hardware.
The SiFive P870 supports "RVA23"[1] which provides more or less feature-parity with ARM. It also supports vector cryptography, which is optional in RVA23. (RVA23 is the successor to last year's RVA22, making Vector and a couple smaller extensions mandatory)
RVA23 + vector crypto also lines up with what Google has announced is likely to be their baseline for RISC-V - based Android handsets in the future.[2]
[1] <https://github.com/riscv/riscv-profiles/blob/main/rva23-prof...>
What appears to be happening is that they bought NUVIA, and a bunch of high performance arm64 cores. ARM sued claiming that cores designed under NUVIA'S Arm Architecture License can't be transferred to Qualcomm's Arm Architecture License. It seems they are trying to slap a riscv decoder with a minimum amount of work. So that means no C extension because the arm64 frontend was designed around aligned 32bit instructions, and tons of custom instructions that match how AArch64 gets passable density, like equivalents to ldp/stp.
So what you're seeing here is a success of RISC-V's model to keep consistency without an incredibly hierarchical body mandating consistency. Qualcomm is more than welcome to create that core and it's numerous extensions. The community is also welcome to not support it by default, and so Qualcomm will probably actually go and complete their design and collaborate.
Was this mentioned anywhere or just a guess?
>It seems they are trying to slap a riscv decoder with a minimum amount of work.
This is interesting, and if true I would just add that they would likely sell both cores – one RISC-V and one ARM – at least in the short term. Proof is that they have already taped out and anounced thier NUVIA based Snapdragon X desktop/laptop SoC. This would further incentivise compatibilty between ARM and RISC-V as this wouldn't be a one time port like Apple did with PA Semi. I would also add that this is all speculation.
Yeah, they're burying it a bit, but it's on their third slide.
> RV64G + 32-bit instructions for code size is best in class
> • More ld/st addressing modes
> • Ld/st pair
> • Conditional immediate branches
> • Move pair
Those are all custom extensions to Qualcomm's design, and suspiciously all how AArch64 gets passable code density.
> This is interesting, and if true I would just add that they would likely sell both cores – one RISC-V and one ARM – at least in the short term. Proof is that they have already taped out and anounced thier NUVIA based Snapdragon X desktop/laptop SoC.
Them announcing that is what made ARM sue them. That core as an AArch64 design is going to be stuck in the courts for longer than it'll be relevant. This work by Qualcomm seems to be a combo of trying to get anything out of the NUVIA acquisition, and maybe having a stick when negotiating a settlement with ARM ("we'll make your existential threat more real").
I'm reffering to this anouncement made October 10.[0]
>2024 will be an inflection point for the PC industry, and Snapdragon X compute platforms will deliver next-level performance, AI, connectivity and battery life.
So it seems that they will just push this out and deal with the consequences later.
[0]https://www.qualcomm.com/news/onq/2023/10/introducing-our-ne...
This doesn’t make sense to me. I thought that one of the selling points of the C extension was that it was a minimal amount of work to add to the decoder.
They claim Qualcomm's arguments are based on flawed statements.
SiFive also argues that Qualcomm's proposal is harmful for microarchitectures of all sizes, including Qualcomm's intended design.
0. https://lists.riscv.org/g/tech-profiles/topic/slides_on_reta...
Some k230 are shipping this month to developers.
An explosion of Vector 1.0 chips and boards is imminent.
0. https://twitter.com/sophgotech/status/1713852202970935334
While I think "information wants to be free" and that stopping CPU collab is like trying to stop encryption in the 90s, I can't help but wonder that RISC-V neuters attempts to lock China out of military AI advancements.
Then again, don't they have x86 factories there? If China has a factory, doesn't the government pretty brazenly learn trade secrets from it?
---
edit to Nickik, since HN has throttled me for more than an hour. Gotta get around the system somehow, since the goddamned rules are so opaque. I'll move this to a comment when I am in the site's good graces again.
---
Highlight where I suggested any of that.
I even said that actually attempting to curtail open development, even if desired, is futile.
I was only wondering if there is validity in suspecting that China will explicitly finance/encourage RISC-V development in order to dodge EU/US sanctions of processor technology.
So all your sarcasm was for nothing.
Are Chinese companies jumping on board this existing effort to avoid sanctions? Almost certainly.
Other than that, don't care much what US congressmen think, since I'm not American, but attempting to close the door on RISC-V will not only be harmful to the broad industry (not just China) but will also be ineffective.
So to be clear, I don't support cramping RISC-V development.