Grace Hopper, Nvidia's Halfway APU
chipsandcheese.com
chipsandcheese.com
Maybe that’s all far enough afield to make the current state of things irrelevant?
Currently rendering and local GPGPU compute is Nvidia dominated and I don’t see AMD competently going after the market segments.
Believe it or not, we've actually been grappling with this scenario for almost a decade at this point. Originally the answer was to unite hardware manufacturers around a common featureset that could compete with (albeit not replace) CUDA. Khronos was prepared to elevate OpenCL to an industry standard, but Apple pulled their support for it and let the industry collapse into proprietary competition again. I bet they're kicking themselves over that one, if they still hold a stronger grudge against Nvidia than Khronos at least.
So - logically, there's actually a one-size-fits-all solution for this problem. It was even going to get managed by the same people handling Vulkan. The problem was corporate greed and shortsighted investment that let OpenCL languish while CUDA was under active heavy development.
> BTW, training is such a high cost that it seems like a major motive for the customer to reduce costs there.
Eh, that's kinda like saying "app development is so expensive that consumers will eventually care". Consumers just buy the end product; they are never exposed to building the software or concerned with the cost of the development. This is especially true with businesses like OpenAI that just give you free access to a decent LLM (or Apple and their "it's free for now" mentality).
Most will probably use something like Llama as base.
Or, in the small business case (mind you, “long term” for tech reaching small businesses is looooong), these businesses again need much smaller models because a) they don’t need a model well versed in Shakespeare and multi variable calculus, and b) they want inference to be as low cost as possible.
These are just scenarios off the top of my head. The broader point is that a dramatic drop in training cost is a wildcard whose effects are really hard to predict.
But I don't know what "long term" is exactly, and have no idea how to time this thing. Besides, I'd bet the sibling evoking the Jevon's paradox is correct.
If one of those scenarios happens, maybe Nvidia can pivot, or if we see analog take over, we could see something really bizarre like a dark horse like Seagate taking over by pivoting from SSDs, just because their manufacturing pipeline is more compatible.
How do we get from here to there, cause I want to get there so bad.
However, I think we need AI beyond current LLMs to really take us there. I'm not saying LLMs can't get us there, we don't know, just beyond what we have. We need AI that we can trust with real tasks IRL.
It is kind of an interesting thought though. A big wall of SSD is a fabulous amount of storage. and maybe a clever read only architecture, would be cheaper than SSD. and a clever data structure for shared high order bits, maybe, maybe there is potential for some device to look up matrix multiply results, or close approximations that could be cheaply refined.
Right now, I doubt it. But big static cache it is a kind of interesting idea to kick around Saturday afternoon.
You are reading the GP the wrong way around.
You store partial results exactly because you can't store computation. Computation is perishable¹, you either use it or lose it. And one way to use it is to create partial results you can save for later.
1 - Well, partially so. Hardware utilization is perishable, but computation also consumes inputs (mostly energy) that aren't. How much it perishable depends on the ratio of those two costs, and your mobile phone has a completely different outlook from a supercomputer.
Shard that across the planet and you'd have a global cache for calculations. Or a lookup for every possible AI prompt and its results.
The flip side of this is if the GPU can access the main system memory then I could see this being useful for loading big models with much more efficient "offloading" of layers. Even though bandwidth between GPU->LPDDR5 is going to be slow, it's still faster than what traditional PCI-E would allow.
The caveat here is that I imagine these machines are $$$ and enterprise only. If something like this was brought to the consumer market though I think it would be very enticing.
(If anybody from AMD is reading this, I feel like an architecture like this would be awesome to have. I would love to run Llama 3.1 405b at home and today I see zero path towards doing that for any "reasonable" amount of money (<$10k?).)
Edit: It's at the bottom of the article. These are designed to be meshed together via NVLink into one big cluster.
Makes sense. I'm really curious how the system RAM would be used in LLM training scenarios, or if these boxes are going to be used for totally different tasks that I have little context into.
- Branch prediction and speculative execution - Out of order execution - Massive physical register files and register renaming - Cache predictors - and many more I'm sure.
Speculative execution is the big one for me, just because of the information leakage possible through it. It's there because you'd have to pause fetching new instructions until the result of a conditional branch is known, which has knock-on effects to instruction scheduling... But how big are these effects? Do some certain combinations of features supercharge or work against each other?
I'm sure there's people looking at such things inside Intel and AMD, but it doesn't seem like there's much out there for public consumption.
Look at how atrocious the CPUs were in the PS4/Xbone generation for an example of this.
I don't know if this was market savvy or a footshoot that made their ecosystem weaker.
In principle, Nvidia could have made chipsets for Intel's newer platforms where the southbridge connects to the CPU over what is essentially four lanes of PCIe, but Intel locked out third parties from that market. But there wasn't much room for Nvidia to provide any significant advantages over Intel's own chipsets, except perhaps by undercutting some of Intel's product segmentation.
(On the AMD side, the DRAM controller was on the CPU starting in 2003, but there was still a separate northbridge for providing AGP/PCIe, with a relatively high-speed HyperTransport link to the CPU. AMD dropped HT starting with their APUs in 2011 and the rest of the desktop processors starting with the introduction of the Ryzen family.)
AFAIR the contentious point was that Nvidia had a license to the bus for P6 arch (by virtue of Xbox) but did not have a license for the P4 bus.
AMD was also more than happy to have NVDA build chipsets for Hammer/etc especially due to them not having a video core... -at the time-.
Once the AMD/ATI merger started, that was the real writing on the wall.
Then Dolby cancelled the license. To this day you still need very fancy sound cards or exotic motherboards to be able to output good surround sound to a large number of av receivers. There are some open DTS standards that Linux can do too, dunno about windows/Mac.
But it just felt like we slid so far down, that Dolby went & made everything so much worse.
(Media software can do Dolby pass-through to let the high quality sound files through, yes. But this means you can't do any effect processing, like audio normalization/compression for example. And if you are playing games your amp may be getting only basic low quality surround surround, not the good many channel stuff.)
I don't know if games are smart enough to use this?
It also feels like a very low bar. It's not awful bitrate for 6 channels but neither is it great. It's not a pitiful number of channels but again neither is it great.
Last & most crucially, just because one piece of software can emit ac3 doesn't make it particularly useful for a system. I should be able to have multiple different apps doing surround sound, sending notifications to back channels or panning sounds as I prefer. Yes ffmpeg can encode 5.1 media audio to an AVR but that doesn't really substitute for an actual surround system.
This is more a software problem, now that the 5.1 AC3 patents are expired. And there have been some stacks in the past where this worked on Linux for example. But it seems like modern hardware (with a Sound Open Firmware) has changed a bit and PipeWire needs to come up with a new way of doing ac3/a52 encoding. https://gitlab.freedesktop.org/pipewire/pipewire/-/issues/32...
That was a long time ago. It is now 2024.
Do we still need that today? For modern AVRs we have HDMI, with 8 channels worth of up to 24bit 192kHz lossless digital audio baked in.
For old AVRs with multichannel analog inputs, motherboards with 6 or 8 channels of built-in audio are still common-enough, as are separate sound cards with similar functionality.
What's the advantage of realtime AC3 encoding today, do you suppose?
FWIW for those unaware like me, 64k refers to 64kB pages, in contrast to the typical 4kB.
[1]: https://www.phoronix.com/review/aarch64-64k-kernel-perf
Nvidia’s Grace Hopper isn’t quite that (it’s primarily a GPU with a bit of CPU sprinkled in), hence “halfway” I guess.
It’s interesting to me that they’ve settled on using standard Neoverse cores, when almost everything else is custom designed and tuned for the expected workloads.
So Nvidia has given up on designing CPU cores, already for some years.
The Carmel core had a performance similar to Cortex-A75, even if it was launched by the time when Cortex-A76 was already available. Moreover, Carmel had very low clock frequencies, which diminished its performance even more. Like also Qualcomm or Samsung, Nvidia has not been able to keep up with the Arm Holdings design teams. (Now Qualcomm is back in the CPU design business only because they have acquired Nuvia.)
I wonder if AMD could license the IBM Telum cache implementation where one core complex could offer unused cache lines to other cores, increasing overall occupancy.
Would be quite neat, even if cross-complex bandwidth and latency is not awesome, it still should be better than hitting DRAM.
Can it run vi?
That might make things much simpler for people who write kernel, drivers and video games.
The history of CPU and GPU prevented that, it was always more profitable for CPU and GPU vendors to sell them separately.
Having 2 specialized chips makes more sense because it's flexible, but since frequencies are stagnating, having more cores make sense, and AI means massively parallel things are not only for graphics.
Smartphones are much modern in that regard. Nobody upgrades their GPU or CPU anymore, might as well have a single, soldered product that last a long time instead.
That may not be the end of building your own computer, but I just hope it will make things simpler and in a smaller package.
> I suspect Grace has a very aggressive prefetcher willing to queue up a ton of outstanding requests from a single core.
Without knowing whether the operations have been performed on FP32 or on FP16 or on another data type, all the numbers written on that page are meaningless.
But this bullshit with Jensen signing girls’ breasts like he’s Robert Plant and telling young people to learn prompt engineering instead of C++ and generally pulling a pump and dump shamelessly while wearing a leather jacket?
Fuck that: if LLMs could write cuDNN-caliber kernels that’s how you would do it.
It’s ok in my book to live the rockstar life for the 15 minutes until someone other than Lisa Su ships an FMA unit.
The 3T cap and the forward PE and the market manipulation and the dated signature apparel are still cringe and if I had the capital and trading facility to LEAP DOOM the stock? I’d want as much as there is.
The fact that your CPU sucks ass just proves this isn’t about real competition just now.
At Wendy’s I get a burger that’s a little smaller every year.
On this I get Enron but smoothed over by Dustin’s OpenPhilanthropy lobbyism.
I’ll take the burger.
edit:
tinygrad IS brat.
YC is old and quite weird.
pytorch but it's minimal so it's not
I was 25 when I apologized for trolling too much on HN, and frankly I’ve posted worse comments since: it’s a hazard of posting to a noteworthy and highly scrutinized community under one’s own name over decades.
I’d like to renew the apology for the low-quality, low-value comments that have happened since. I answer to the community on that.
To you specifically, I’ll answer in the way you imply to anyone with the minerals to grow up online under their trivially permanent handle.
My job opportunities and livelihood move up and down with the climate on my attitudes in this forum but I never adopted a pseudonym.
In spite of your early join date which I respect in general as a default I remain perplexed at what you’ve wagered to the tune of authenticity.
You joined early, I’ve been around even longer.
You can find a penitent post from me about an aspiration of higher quality participation, I don’t have automation set up to cherry-pick your comments in under a minute.
My username is my real name, my profile includes further PII. Your account looks familiar but if anyone recognizes it on sight it’s a regime that post-dates @pg handing the steering wheel to Altman in a “Battery Club” sort of way.
With all the respect to a fellow community member possible, and it’s not much, kindly fuck yourself with something sharp.
My "opposition research" was entirely two clicks, profile (see account age), Submissions (see oldest).
As for pseudonym's, I've been online since Usenet and have never once felt the need to advertise on the new fangled web (1.0, 2.0, or 3), handles were good enough for Ham Radio, and TWKM - Those Who Know Me Know Who I Am (and it's not at all that interesting unless you like yarns about onions on belts and all that jazz).
I want to apologize sincerely for my recent comments, particularly my last response to you. Upon reflection, I realize that my words were hurtful, disrespectful, and completely inappropriate, especially given the light-hearted nature of your previous comment. I am truly sorry for any offense or harm I may have caused.
Your comment was clearly intended as a friendly jest, and I regret that I responded with such hostility. There is no excuse for my behavior, and I am committed to learning from this mistake and ensuring it does not happen again.
I also want to address my earlier comments in this thread. I now understand that my attempts to justify my past behavior and dismiss genuine concerns came across as defensive and disrespectful. Instead of taking responsibility for my actions, I tried to deflect and downplay their impact, which only served to escalate the situation.
I value this community and the opportunity it provides for open dialogue and growth. I understand that my actions have consequences, and I am determined to be more mindful, respectful, and considerate in my future interactions. I promise to strive for higher quality participation and to treat all members of this community with the kindness and respect they deserve.
Once again, I am truly sorry for my offensive remarks and any harm they may have caused. I appreciate the understanding and patience you and the community have shown, and I hope that my future actions will reflect my commitment to change and help rebuild any trust that may have been lost.
Again, no drama - my sincere apologies for inadvertently poking an old issue, there was no intent to be hurtful on my part.
I have a thick skin, I'm Australian, we're frostily polite to those we despise and call you names if we like you - it can be offputting to some. :)
You’ve been a good sport legend.