Intel announces Arc B-series "Battlemage" discrete graphics with Linux support
phoronix.com
phoronix.com
At the very least, it's nice to have some decent BUDGET cards now. The ~$200 segment has been totally dead for years. I have a feeling Intel is losing a fair chunk of $ on each card though, just to enter the market.
https://christianjmills.com/posts/intel-pytorch-extension-tu...
https://chsasank.com/intel-arc-gpu-driver-oneapi-installatio...
I quite enjoyed using the CUDA to OneAPI migration tool. Was a bit dodgy at times, but for the most part helped me move a lot of stuff out of the NVIDIA walled garden.
Uh...yes they do? I mean, look at all the VC startups these days like Uber/Doordash/Rivian/WeWork/Carvana/etc. Some of those have been bleeding hundreds of millions per quarter for years and keep getting more pumped in over and over.
But even when it's not for insane VC reasons, companies burn money to get market share all the time. It's a textbook strategy. They just don't do it forever. Eventually, the shareholders can get impatient and decide to bail on the idea.
1920x1080 is 1080p.
It doesn't make a whole lot of sense, but that's how it is.
[EDIT] I mean, of course, 1080p's also not typically called that, yet another resolution is, but labeling 1440p 2k is especially far off.
> (the p is progressive, vs i for interlaced)
> 4k, 2k refer to the number of columns of pixels
> 2560×1440 is both 2.5k and 1440p, and 3840×2160 is both 4k and 2160p.
These parts I did not misunderstand.
> and those terms come from cinema and visual effects (and originally means 4096 and 2048 pixels wide)
OK that part I didn't know, or at least had forgotten—which are effectively the same thing, either way.
> 1920×1080 is both 2k and 1080p
Wikipedia suggests that in this particular case (unlike with 4k) application of "2k" to resolutions other than the original cinema resolution (2048x1080) is unusual; moreover, I was responding to a commenter's usage of "2k" as synonymous with "1440p", which seemed especially odd to me.
“2K” is used to denote WQHD often enough, whereas 1080p is usually called that, if not “FHD”.
“2K” being used to denote resolutions lower than WQHD is really only a thing for the 2048 cinema resolutions, not for FHD.
Assuming they retail at prices Intel are suggesting in the press releases, you maybe here save 40-50 bucks over an ~equivalent NVidia 4060.
I would also argue like others here that with tech like frame gen, DLSS etc, even the cheapest discrete NVidia 40xx parts are arguably 1440p optimized now, it doesn't even need to be said in their marketing materials. Im not as familiar with AMD's range right now, but I suspect virtually every discrete graphics card they sell is "2k optmized" by the standard Intel used here, and also doesn't really warrant explicit mention.
Time marches on, and I become ever more separated from gaming PC enthusiasts.
Instead, I bought a good 27" 1440p monitor, and you know what? I am not the discerning connoisseur of pixels that I thought I was. Honestly, it's fine.
I will hold out with this setup until I can get a 8k 144hz monitor and a gpu to drive it for a reasonable price. I expect that will take another decade or so.
At 4K, it's like having 4 21" 1080p monitors. Haven't maximized or minimized a window in years. The sprawl is real.
Get a 5K 27".
No fractional scaling, same real estate, much better picture.
In other words, the current tech just isn’t quite there yet, or not cheap enough.
It’s a great way to have my cake and eat it too.
You say PC gamers at the start of your comment and gaming PC enthusiasts at the end. These groups are not the same and I'd say the latter is largely doing ultrawide, 4k monitor or even 4k TV.
According to steam, 56% are on 1080p, 20% on 1440p and 4% on 2160p.
So gamers as a whole are still settled on 1080p, actually. Not everyone is rich.
It does have HDMI-CEC, so I haven't even used the remote control in several years.
If you’re ok with the resolution, then the only downside is significant power consumption and lack of HDR support.
Prove to me those aren't synonyms.
Ouch, had something similar happen to me before when I bought a VR headset and had to return it. Wishing you the best on your job search!
Could just get a 3060 and a nice 1440p monitor.
Even with this monitor, I'm barely able to run it with my (expensive, though older) graphics card, and the screen alarmingly flashes whenever I change any settings. It's stable, but this is not a simple plug-and-play configuration (mine requires two DP cables and fiddling with the menu + NVIDIA control panel).
https://linustechtips.com/topic/729232-guide-to-display-cabl... is a calculator for bandwidth 4K@144 HDR is ~40 Gigabit/s. You can do better with compression, but I find Nvidia cards have an issue with compression enabled.
I don't think it's an issue until you notice. I only noticed because I toggle HDR for some games and at 1440p240hz, the difference is just enough to not need DSC
Reddit also seems to have some people who have managed to get 144 with FreeSync, but I've only managed 120.
Funnily enough while I was typing this Netflix caused both my monitors to blackscreen (some sort of NVIDIA reset I think) and then come back. It's not totally stable!
It works up until too many pixels change, basically.
1440p hits a popular balance where it’s more pixels than 1080p but not so absurdly expensive or power hungry.
Eventually 4K might be reasonably affordable, but we’ll settle at 1440p for a while in the meantime like we did at 1080p (which is still plenty popular too).
Like I said, it's on the cusp of invisible pixels.
I've not tried but I've heard that a butter-smooth 90, 120, or 300 FPS frame rate (that is also synchronized with the display) is really wonderful in many such games, and once you experience that you can't go back. On less powerful systems it then requires making a tradeoff with rendering quality and resolution.
Tbh now that I think about it I only really need resolution for general usage. For gaming I'm running everything but textures on low with min or max FOV depending on the game so it's not exactly aesthetic anyway. I more so need physical screen size so the heads are physically larger without shoving my face in it and refresh rate.
Most people don't have enough disposable income to make spending that extra amount a reasonable tradeoff (and continuing to spend on upgrades to keep up with their monitor on new games).
It's doable, the tech is there. But the cost is WAY too high compared to what you get from it in the end.
I run the second monitor off the IGPU so it doesn't even tax the main GPU.
...which will be most likely in better condition than anything that was used for, let's say, gaming...
Untouched video (star wars 8) 4k HDR (60Mb/s) to 1080p at 28fps
No SR-IOV.
No issues with linux. The server did not like the a310, but that is because it is an old dell t430 and it is unsupported hardware. The only thing I had to do was to tweak the fan curve so that it stopped going full tilt.
I guess before you used Nvidia because out of the box support for AMD has existed for ages.
Video games
These ML AI Macbook people are legit insane.
Desktops and gaming is ugly and complex to them (because lego is hard and macbook look nice unga bunga), yet it is a mass market Intel wants to move in on.
People here complain because Intel is not making a cheap GPU to "make AI" on when that's a market of maybe 1000 people.
This Intel card is perfect for an esports gaming machine running CS2, Valorant, Rocket Leauge and casual or older games like The Sims, GoG games etc. Market of 1 million + right there, CS2 alone is 1mil people playing everyday. Not people grinding leetcode on their macs. Every real developer has a desktop, epyc cpu, giga ram and a nice GPU for downtime and run a real OS like Linux or even Windows (yes majority of devs run Windows)
>market of maybe 1000 people
The market of people interested in local ai inference is in the millions. If it's cheap enough the data center market is at least 10 million.
/s
Intel has only had discrete GPUs on the market for 2 years. I guess that is a plural number of years, but only barely.
Both groups have a high autism %
We love to be "technically correct" and we often are. So we get frustrated when people claim things that are wrong.
B770 was rumoured to match the 16 GB of the A770 (and to be the top end offering for Battlemage) but it is said to not have even been taped out yet with rumour it may end up having been cancelled completely.
I.e. don't hold your breath for anything consumer from Intel this generation better for AI than tha A770 you could have bought 2 years ago. Even if something slightly better is coming at all there is no hint it will be soon.
Hm, i wouldn't consider 200$ low end.
The Intel cards are getting more interesting for me as I'm questioning my continued use of macOS. Intels focus on Linux support makes their options really interesting, though I don't see a need for something as powerful as these new cards.
The Radeon 780M on Ryzen APUs can power 1080P gaming, and output to 3-4 4K/8K displays (the latest Intel Xe iGPUs are about on par, but generally pricier). A Ryzen 5700G goes for about $150 (or a 5600G for closer to $100), or you can get entire Ryzen 7840HS minipcs w/ 32GB RAM for $400-500).
Or we can keep asking high computers questions about programming.
I agree ML is about to hit (or has likely already hit) some serious constraints compared to breathless predictions of two years ago. I don't think there's anything equivalent to the AI winter on the horizon, though—LLMs even operated by people who have no clue how the underlying mechanism functions are still far more empowered than anything like the primitives of the 80s enabled.
Nvidia had a revenue of $27billion in 2023 - that's about $160 per person per year [0] for every working age person in the USA. And it's predicted to more than double in 2024. If you reduce that to office workers (you know, the people who might actually get some benefit, as no AI is going to milk a cow or serve you starbucks) that's more like $1450/year. Or again more than double that for 2024.
How much value add is the current set of AI products going to give us? It's still mostly promise too.
Sure, like most bubbles there'll probably still be some winners, but there's no way the current market as a whole is sustainable.
The only way the "maximal AI" dream income is actually going to happen is if they functionally replace a significant proportion of the working population completely. And that probably would have large enough impacts to society that things like "Dollars In A Bank" or similar may not be so important.
[0] Using the stat of "169.8 million people worked at some point in 2022" https://www.bls.gov/news.release/pdf/work.pdf
[1] 18.5 million office workers according to https://www.bls.gov/news.release/ocwage.nr0.htm
Seems like a lot of CV solutions have seen fairly steady but small incremental advances over the past 10-15 years, quite unrelated to the current AI hype.
We've been through multiple AI Winters, as a new technique is developed, it does increase the capabilities. Just not as much as the hype suggested.
To say there won't be a bust implies this boom will last forever, into whatever singularity that implies.
For example ?
(besides deep fakes)
Used them in the garden while weeding and in the garden store while planning what to plant, in both cases to identify plants by image and tell me about them — though I'd say image capable AI are no longer mere "large language models".
Used ChatGPT while shopping to help me locate products in the store I was in, when I couldn't find just by wandering the isles, by uploading a photo of the aisle I happened to be in at the point I gave up.
> Nvidia had a revenue of $27billion in 2023 - that's about $160 per person per year [0] for every working age person in the USA
As a non-American, I'd like to point out we also earn money.
> as no AI is going to milk a cow or serve you starbucks
Cows have been getting the robots for a while now, here's a recent article: https://modernfarmer.com/2023/05/for-years-farmers-milked-co...
Robots serve coffee as well as the office parts of the coffee business: https://www.techopedia.com/ai-coffee-makers-robot-baristas-a...
Some of the malls around here have food courts where robots bring out the meals. I assume they're no more sophisticated than robot vacuum cleaners, but they get the job done.
Transformer models seem to be generally pretty good at high-level robot control, though IIRC a different architecture is needed down at the level of actuators and stepper motors.
And I know restricting it to the US is a simplification, but so is restricting it to Nvidia, it's just to give a ballpark back-of-the-envelope "does this even make sense?" level calculation. And that's what I'm failing to see.
Nonetheless, Starbucks does not use these machines, and I don't see any reason that AI, on its current trajectory, will change that calculation any time soon.
They could serve us a plate of shit and we'd debate if pepper or salt is better to complement it
I mean, Yudkowsky has basically spent the last decade screaming into the void about how AI will with high probability literally kill everyone, and even people like me who think that danger is much less likely still look at the industrial revolution and how slow we were to react to the harms of climate change and think "speed-running another one of these may be unwise, we should probably be careful".
I'm still not convinced about that. All the """studies""" show 30-60% boost in productivity but clearly this doesn't translate to anything meaningful in real life because no industry laid off 30-60% of their workforce and no industry progressed anywhere close to 30% since chat gpt was released.
It's been released a whole 24 months ago, remember the talks about freeing us from work and curing cancer... Even investments funds which are the biggest suckers for anything profitable are more and more doubtful
The services that provide serious productivity boosts aren't being heavily used or marketed yet. They:
1. Attempt to do narrow tasks with high proficiency. 2. Replace specific job titles. 3. Are high-value enough to be slower to engage in layoffs.
What I'm suggesting is not a fungible solution, but it is one that will be highly profitable and productive.
I really wasn't interested in computer hardware anymore (they are fast enough!) until I discovered the world of running LLMs and other AI locally. Now I actually care about computer hardware again. It is weird, I wouldn't have even opened this HN thread a year ago.
It has zero cost, hardware is already there. I'm not captive to some remote company.
I can fiddle and integrate with other home sensors / automation as I want.
Hardware side, I just have a beefy server that acts as a router (mellanox card to provider fiber optic and local fiber network), firewall, wifi access point, zigbee coordinator, host to various services, camera video feed ingestion and processing, and so on...
Honestly, Apple seems to be on the right track here. DDR5 is slower than GDDR6, but you can scale the amount of RAM far higher simply by swapping out the density.
People did it with the RTX3070. https://www.tomshardware.com/news/3070-16gb-mod
To have more memory, you have to design a new die with a wider interface. The design+test+masks on leading edge silicon is tens of millions of NRE, and has to be paid well over a year before product launch. No-one is going to do that for a low-priced product with an unknown market.
The savior of home inference is probably going to be AMD's Strix Halo. It's a laptop APU built to be a fairly low end gaming chip, but it has a 256-bit LPDDR5X interface. There are larger LPDDR5X packages available (thanks to the smartphone market), and Strix Halo should be eventually available with 128GB of unified ram, performance probably somewhere around a 4060.
You're right: nobody's doing ML these days. /s
Look at the value of a single website with well over a million users where they publish and run open-weights models on the regular: back in August of 2023, Huggingface's estimated value was $4.5 billion.
Can you even do ML work with a GPU not compatible with CUDA? (genuine question)
A quick search showed me the equivalence to CUDA in the Intel world is oneAPI, but in practice, are the major Python libraries used for ML compatible with oneAPI? (Was also gonna ask if oneAPI can run inside Docker but apparently it does [1])
Vulkan is especially appealing because you don't need any special GPGPU drivers and it runs on any card which supports Vulkan.
This is a graphics card.
tl;dr GPU's need to transition from being add-in cards to being a sibling motherboard. A sisterboard? Not a daughter board.
Intel and AMD internal GPUs can use normal computer RAM. But they are slower for that reason and many others.
It's one of the reasons why ARM Macbooks get great performance/watt, memory being even "closer" than mainboard soldered RAM so getting more of those benefits, though naturally less flexibility.
[0] https://store.steampowered.com/hwsurvey/videocard/ - 0.19% share
-.-
I feel like _anyone_ who can pump out GPU's with 24GB+ of memory that are capable to use for py-stuff would benefit greatly.
Even if it's not as performant as the NVIDIA options - just to be able to get the models to run, at whatever speed.
They would fly off the shelves.
At some point they should make a stand, that's the whole meta-topic of this thread.
If you get the base 16GB mini, it will have more or less the same VRAM but way worse performance than an Arc.
If you already have a PC, it makes sense to go for the cheapest 12GB card instead of a base mac mini.
I don't know where this idea is coming from, although it's all over these threads.
For context, I write a local LLM inference engine and have 0 idea why this would shift anyone's purchase intent. The models big enough to need more than 12GB VRAM are also slow enough on consumer GPUs that they'd be absurd to run. Like less than 2 tkns/s. And I have 64 GB of M2 Max VRAM and a 24 GB 3090ti.
You'd need hundreds of thousands of units to really make much of a difference.
Was it? Their GPUs sales were insignificantly low so I doubt that had a huge effect on their net income.
It's going to be a HomePod/AppleTV/Echo/Google Home -style box you set up in a corner and forget about it.
Then your devices in the ecosystem can offload some LLM tasks to that local system for inference, without having to do everything on-device.
I mean, everything could have been already working that way for a lot of years right? One big shared compute box in your house and everything else is a dumb screen? But few people roll that way, even nerds, so I don't see that becoming a thing for offloaded AI compute.
I also think that the future of consumer AI is going to be models trained/refined on your own data and habits, not just a box in your basement running stock ollama models. So I have some latency/bandwidth/storage/privacy questions when it comes to wirelessly and transparently offloading it to a magic AI box that sits next to my wireless router or w/e, versus running those same tasks on-device. To say nothing of consumer appetite for AI stuff that only works (or only works best) when you're on your home network.
Both are currently used as the hub for HomeKit devices. Making the ATV into a "magic AI Box" won't need anything else except "just" upgrading the CPU from A-series to M-series. Actually the A18 Pro would be enough, it's already used for local inference on the iPhone 16 Pro.
Would it though? How many people are running inference at home?
I don't know how to quantify it, but it certainly seems like a lot of people are buying consumer nVidia GPUs for compute and the relatively paltry amounts of RAM on those cards seems to be the number one complaint.So I would say that Intel's potential market is "everybody who is currently buying nVidia GPUs for compute."
nVidia's stingy consumer RAM choices also seem to be a fairly transparent ploy to create a protective moat around their insanely-high-profit-margin datacenter GPUs. So that just seems like kind of an obvious thing for Intel or AMD to consider tackling.
(Although, it has to be said, a lot of commenters have pointed out that it's not as easy as just slapping more RAM chips onto the GPU boards; you need wider data busses as well etc.)
Specifically, waiting ~2 seconds vs ~20 for a code snippet is much more detrimental to my productivity than the time difference would suggest. In ~2 seconds I don't get distracted, in ~20 seconds my mind starts wandering and then I have to spend time refocusing.
Make a GPU that is 50% slower than a 2 generations older mid-range GPU (in tokens/s) but on bigger models and I would gladly shell out 1000+$.
So much so that I am considering getting a 5090 if nVdia actually fixes the connector mess they made with 4090s or even a used v100.
I think that's what GP was alluding to.
Running a specialist model makes more sense on small devices.
Like the 4060 Ti would have been a nice fit if it hadn't been for the narrow memory bus, which makes it slower than my 2080 Ti for LLM inference.
A more expensive card has the downside of not being cheap enough to justify idling in my server, and my gaming card is at times busy gaming.
Well informed gamers know Intel's discrete GPU is hanging by a thread, so they're not hoping on that bandwagon.
Too small for ML.
The only people really happy seem to be the ones buying it for transcoding and I can't imagine there is a huge market of people going "I need to go buy a card for AV1 encoding".
They do well compared to AMD/Nvidia at that price point.
Is it a market worth chasing at all?
Doubt.
Had this been available a few weeks ago I would have gone through the pain of early adoption. Sadly it wasn't just an upgrade build for me, so I didn't have the luxury of waiting.
What do you mean by this - I assume you mean too small for SoTA LLMs? There are many ML applications where 12GB is more than enough.
Even w.r.t. LLMs, not everyone requires the latest & biggest LLM models. Some "small", distilled and/or quantized LLMs are perfectly usable with <24GB
Still...tangibly cheaper than even a 2nd hand 3090 so there is perhaps a market for it
Intel's customers are 3rd party Cpu assemblers like Dell & HP. Many corporate bulk buyers only care if 1-2 of the apps they use are supported. The lack of wider support isn't a concern.
Nvidia is trash tier in terms of support and only recently making serious steps to actually support the platform.
AMD went all in nearly a decade ago and it's working pretty well for them. They are mostly caught up to being Intel grade support in the kernel.
Meanwhile, Intel has been doing this since I was in college. I was running the i915 driver in Ubuntu 20 years ago. Sure their chips are super low power stuff, but what you can do with them and the level of software support you get is unmatched. Years before these other vendors were taking the platform seriously Intel was supporting and funding Mesa development.
That said, I wonder if you are using the nouveau driver. The Nvidia driver should not have such issues.
They have been improving because getting CUDA running on the platform in datacenters required some level of cooperation. They eventually realized this had to happen regardless for AI in the datacenter and so eventually they got a kernel module starting to be added. It's not really there yet and it doesn't run the actual driver, but it gets a pathway to booting Nvidia with signed drivers among other things.
The end result is our support forums are still full of Nvidia users. Go to any support forum or Reddit related to Linux and view the comments and press CTRL+F for Nvidia. It's so common we have to just tag/bin these posts and move on because there are literally man years of time being wasted on this.
> AMD on the other hand gives lousy support.
I can open a bug right now on both GitHub and the LKMS directly with AMD engineers. When I have a question about ROCm I can ask them and get a quick response. These guys are in Reddit threads talking to ROCm users and giving them guidance on how to modify their consumer grade Radeon cards to use the datacenter grade drivers that aren't adapted to consumer facing cards yet.
The one thing we for sure agree on is that Intel doesn't have any of this BS. Their stuff just works and their pipeline for getting anything fixed is also well optimized.
As for opening a bug on GitHub, there is a difference between them responding and them fixing things. I have seen many people report bugs without them fixing the issues. There is one prominent issue for parity between direct3d q2 and vulkan that was needed by emulators that they outright refused to fix:
https://github.com/GPUOpen-Drivers/AMDVLK/issues/108
With Intel, the last time I opened a bug with them, they fixed the issue I reported within 24 hours.
Ubuntu 24.04 couldn't even boot to a tty with the Nvidia Quadro thing that came with this major-brand PC workstation, still under warranty.
Why would that matter? You buy one GPU, in a few years you buy another GPU. It's not a life decision.
The game devs are going to spend all their time & effort targetting amd/nvidia. Custom code paths etc.
It's not a one size fits all world. OpenCL etc abstraction are good at covering up differences, but not that good. So if you're the player with <10% market share you're going to have an uphill battle to just be on par.
From my experience they target NVidia and Consoles. AMD might get a look at the code just before release if they notice any big problems.
I'd be surprised if many Gamedevs even pick up the phone for an Intel GPU developer.
In particular, intel just needs to support vfio and it’ll be huge for homelabs.
If Intel's stats are anything to go by, League runs way better than it did on the last generation and it's the only game that has had issues on the last-gen that's still left running on DX9, CS:GO was another notable one but CS2 has launched since and the game has moved to DX12/VK. This was, literally, the biggest issue they had - drivers were also wonky but they seem to have ironed that out as well.
there is a niche for this, though it remains to be seen if it'll be profitable enough for a large org like Intel
I am hoping these are open in such a manner that they can be used in OpenBSD. Right now I avoid all hardware with a Nvidia GPU. That makes for somewhat slim pickings.
If the firmware is acceptable to the OpenBSD folks, then I will happly use these.
I don't think that's a super fair shake? Intel iGPUs have been around for a while if you had a laptop chip or iGPU-enabled desktop chip. They've supported Linux just fine for ages, and will fill any non-3D application you might have.
And Nvidia chips are quite good on Linux nowadays - Wayland has been very usable since the 535-series drivers and nearly flawless since 550. You're right to be apprehensive about proprietary GPU hardware but I think there are plenty of options on the table right now.
For power, it's 190W compared to 4060's 115 W.
EDIT: from [1]: B580 has 21.7 billion transistors at 406 mm² die area, compared to 4060's 18.9 billion and 146 mm². That's a big die.
https://www.techpowerup.com/review/intel-arc-b580-battlemage...
These numbers seem bit more believable
If we use the numbers from the preview:
| |Arc A770|Arc B580|RTX 4060|
|--------|--------|--------|--------|
|Process |N6 |N5 |N5 |
|Die Size|406mm^2 |272mm^2 |159mm^2 |
|Trans. |21.7B |19.6B |18.9B |
|Mem Bus |256 bit |192 bit |128 bit |
|TDP |225W |190W |115W |
|~Perf |90% |110% |100% |
In terms of performance per die area it's a big improvement over A770 but still far behind Nvidia. It's interesting that the transistor density is so much lower than the 4060 despite having the same (or at least similar) process node. Speculating about why that may be:- Nvidia has better layout. - Intel is using higher performance (larger) transistor libraries or layout in order to hit the higher boost frequencies (2800 vs 2460). - Wider bus interface takes up more space. - The B580 may have 1 render slice and 64-bits of memory bus disabled, and they're not including those transistors in the count, but they still take up area.
[0]: https://www.techpowerup.com/review/intel-arc-b580-battlemage...
1. There are dummy transistors included in the design for various manufacturing reasons, and there's no standard as to whether those are included in the numbers.
2. SRAM cells in particular are highly optimized and the amount of SRAM (e.g. cache memory) on the chip versus logic will have a big influence on the transistor density.
He said he's still asking around the engineering team for a more complete explanation but that's the supposition.
It's a real shame, the single slot a380 is a great performance for price light gaming and general use card for small machines.
I cling onto my old hardware to limit ewaste where I can. I still gave up on my old sandybridge machine once it hit about a decade old. Not only would the CPU have trouble keeping up, its mostly only PCIe 2.0. A few had 3.0. You wouldn't get the full potential even out of the cheapest one of these intel cards. If you are putting a GPU in a system like that I can't imagine even buying new. Just get something used off ebay.
Consumer boards and CPUs didn't really support it well until after 2018. I upgraded away from a Zen+ system because it didn't support it.
Another issue is that not every GPU actually supports ReBAR, I'm reasonably certain the Nvidia drivers turn it off for some titles, and pretty much the only vendor that reliably wants ReBAR on at all times is Intel Arc.
I also personally wouldn't say that Sandy Bridge is very usable with a modern GPU without also specifying what kind of CPU or GPU. Or context in how it's being used.
Coffee Lake also didn't really support ReBAR either, also 2018.
Also, I don't think you'll find many mainboards from 2006 supporting it. It may have been standardized in 2006, but a quick online search leads me to think that even on x86 mainboards it didn't become commonly available until at least 2020.
What I'm also interested in is whether they'll require a PCIE 4.0 motherboard or if a 3.0 x16 slot is fine.
Also whether their software properly supports VR now...
I haven't regretted the purchase at all.
The only concern is how well the new Intel drivers work (full support for DX12) with older titles which are continuously being improved (for DX11, 10, and some for 9 others via emulation).
There's likely some deep discounting of Intel cards because of how bad the drivers were at launch and the prices may not stay so low once things are working much better.
IMO you want those frames getting rendered as close to the monitor as possible, and you'd probably have a better time with lower fidelity graphics rendered locally. You'd also get to keep gaming during a network outage.
I've tried game streaming under the best possible conditions (<1ms network latency) and it still feels a little off. Especially shooters and 2D platformers.
Fortunately, having their Linux drivers be (mostly?) open source makes a purchase seem less risky.
These are obviously Windows-specific issues that don't come up at all in Linux, where all that Direct3D headache is taken care of by DXVK. Amusingly a big part of Intel's efforts to improve D3D performance on Windows has been to use DXVK for many titles.
It is in the (minor) sense that I'd rely on Intel for warranty support, driver updates (if closed source), and firmware fixes.
But I agree with your main point that the worst-case downside isn't that big of a deal.
I agree entirely.
My point was that even if Intel disappeared tomorrow, there's a good chance that Linux developer community would take over maintenance of those drivers.
In contrast to, e.g., 10-years-ago nvidia, where IIUC it was very difficult for outsiders to obtain the documentation needed to write proper drivers for their GPUs.
Battlemage is supposed to fix all these architectural issues. EU in Xe2 is now SIMD16 (which is why the number of EUs per Xe2 core is halved from that of Xe1), and they've added all the previously software-emulated instructions, including Execute Indirect, so in theory Battlemage should be in a much better position in game compatibility side of things.
On Linux side of things, lacking sparse residency support in i915 also contributes to game compatibility[1] (though this is now available under Mesa 24). This is something the new xe driver is supposed to fix, but it's still a long way to go until it's actually usable.
I've had a small sound latency issue forever; most visible with YouTube videos, the first half-second of every video is silent.
I picked this card up for about $120 less than the GTX 4060. Wasn't a terrible decision.
Presumably graphics cards optimised for hairdressers and telephone sanitisers?
Other exciting tests will include things like fan control, since that’s still an issue with Arc GPUs.
Should make for a fun blog post.
Very happy with my A770. Godsend for people like me who want plenty VRAM to play with neural nets, but don't have the money for workstation GPUs or massively overpriced Nvidia flagships. Works painlessly with linux, gaming performance is fine, price was the first time I haven't felt fleeced buying a GPU in many years. Not having CUDA does lead to some friction, but I think nVidia's CUDA moat is a temporary situation.
Prolly sit this one out unless they release another SKU with 16G or more ram. But if Intel survives long enough to release Celestial, I'll happily buy one.
My experience on WH40K DT has taught me that upscaling is absolutely vital for a reasonable experience on some games.
This strikes me as a bit of a sad state of affairs. We've moved beyond a Parkinson's law of computational resources –usage by games expands to fill the available resources– to resource usage expanding to fill the available resources on the highest end machines unavailable for less than a few thousand dollars... and then using that to train a model to simulate by upscaling higher quality or performance on lower end machines.
A counterargument would be that this makes high-end experiences available to more people, and while in the individual case, I don't buy that that's where the incentives it creates are driving the entire industry.
To put a finer point on it: at what percentage of budget is too much money being spent on producing assets?
What a time to be alive. Our most advanced technology is used to cheat on homework and play video games.
It used to be that more computational power was desirable because it would allow for developers to more fully realize creative visions that weren't previously possible.
Now, it seems that the goal is simply visual fidelity and asset complexity... and the rest of the experience is not only secondary, but compromised in pursuit of the former.
Thinking back on recent games that felt like something new and painstakingly crafted... they're almost all 2D (or look like it), lean on excellent art/music (and even haptics!) direction, have a well-crafted core gameplay loop or set of systems, and have relatively low actual system requirements (which in turn means they are exceptionally smooth without any AI tricks).
Off the top of my head few years: Hades, Balatro, Animal Well, Cruelty Squad[0], Spelunky, Pizza Tower, Papers Please, etc. Most of these could just as easily have been made a decade ago.
That's not to say we haven't had many games that are gorgeous and fun. But while the latter is necessary and sufficient, the former is neither.
It's just icing: it doesn't matter if the cake tastes like crap.
[0] a mission statement if there ever was one for how much fun something can be while not just being ugly but being actively antagonistic to the senses and any notion of good taste.
That's not _quite_ how temporal upscaling work in practice. It's more of a blend between existing pixels, not generating entire pixels from scratch.
The technique has existed since before ML upscalers became common. It's just turned out that ML is really good at determining how much to blend by each frame, compared to hand written and tweaked per-game heuristics.
---
For some history, DLSS 1 _did_ try and generate pixels entirely from scratch each frame. Needless to say, the quality was crap, and that was after a very expensive and time consuming process to train the model for each individual game (and forget about using it as you develop the game; imagine having to retrain the AI model as you implement the graphics).
DLSS 2 moved to having the model predict blend weights fed into an existing TAAU pipeline, which is much more generalizable and has way better quality.
But surely it's easy enough to compete on video ram - why not load their GPUs to the max with video ram?
And also video encoder cores - Intel has a great video encoder core and these vary little across high end to low end GPUs - so they could make it a standout feature to have, for example, 8 video encoder cores instead of 2.
It's no wonder Nvidia is the king because AMD and Intel just don't seem willing to fight.
Also having different encoding settings for different purposes is desired (e.g. high quality local recording for an edit later while live streaming to different services at the same time [Twitch, Youtube, ...]).
That said, I'm not aware of Intel limiting the number of encoding streams, so I don't know where the number 2 originates.
I think it's the right call since there isn't much competition in GPU industry anyway. Sure, Intel is far behind. But they need to start somewhere in order to break ground.
Strictly speaking strategically, my intuition is that they will learn from this, course correct and then would start making progress.
> "They need to start somewhere in order to break ground"
Intel has big problems and it's not clear they should occupy themselves with this. They should stabilize, and the most plausible way to do that is to cut the weak parts, and get back to what they were good at - performant secure x86_64 CPUs, maybe some new innovative CPUs with low consumption, maybe memory/solid state drives.
That's a very low margin and cyclical market since memory/SSDs are basically commodities. I don't think Intel would have any chance surviving in such a market they just have way to much bloat/R&D spending. Which is not a bad thing as long as you can produce better than products than the competition.
"there is not much money in it"?
WTF?
Intel tried to get into GPU-like products 14 years ago. They promote their consumer Arc GPUs since 2022, and still almost nobody wants them. Is there a datacenter Intel GPU that some business wants?
The big money is flowing NVIDIA's way, and even if Intel can make a GPU, it won't be able to divert a big part of the flow, similarly to AMD.
However, this is going to go on clearance within 6 months. Good for consumers, bad for Intel.
Also keep in mind for any ML task Nvidia has the best ecosystem around. AMD and Intel are both like 5 years behind to be charitable...
How was driver support for their A-series?
They likely won't need to do the same discovery and fixing for B-series as they've already dealt with it.
Not a huge fan of the numbering system they've used. B > A doesn't parse as easily as 5xxx > 4xxx to me.
I kinda like the idea of Intel.
IIRC that was one of the original goals of geohot's tinybox project, though I'm not sure exactly where that evolved
Their dedication to Linux Support, combined with their good pricing makes this a potential buy for me in future versions. To be frank, I won't be replacing my 7900 XTX with this. Intel needs to provide more raw power in their cards and third parties need to improve their software support before this captures my business.
ah well. pretty sure it'll do for my needs.
Based on scaling by XMX/engine clock napkin math, the B580 should have 230 FP16 TFLOPS and 456 GB/s MBW theoretical. At similar efficiency to LNL Xe2, that should be about pp512 ~4700 t/s and tg128 ~77 t/s for a 7B class model. This would be about 75% of a 3090 for pp and 50% for tg (and of course, 50% of memory). For $250, that's not too bad.
I do want to note a couple things from my poking around. The IPEX-LLM [1] was very responsive, and was able to address an issue I had w/ llama.cpp within days. They are doing weekly update releases, so that's great. The IPEX stands for Intel Extension for PyTorch [2] and it is a mostly drop-in for PyTorch: "Intel® Extension for PyTorch* extends PyTorch* with up-to-date features optimizations for an extra performance boost on Intel hardware. Optimizations take advantage of Intel® Advanced Vector Extensions 512 (Intel® AVX-512) Vector Neural Network Instructions (VNNI) and Intel® Advanced Matrix Extensions (Intel® AMX) on Intel CPUs as well as Intel Xe Matrix Extensions (XMX) AI engines on Intel discrete GPUs. Moreover, Intel® Extension for PyTorch* provides easy GPU acceleration for Intel discrete GPUs through the PyTorch* xpu device."
All of this depends on Intel oneAPI Base Kit [3] which has easy Linux (and presumably Windows) support. I am normally an AUR guy on my Arch Linux workstation, but those are basically broken and I had much more success installing oneAPI Base Kit (w/o issues) directly in Arch Linux. Sadly, this is also where there are issues some of the code is either dependent on older versions of oneAPI Base Kit that are no longer available (vLLM requires oneAPI Base Toolkit 2024.1 - this is not available for download from the Intel site anymore) or in dependency hell (GPU whisper simply will not work, ipex-llm[xpu] has internal conflicts from the get go), so it's not all sunshine. On average, ROCm w/ RNDA3 is much more mature (while not always the fastest, most basic things do just work now).
[1] https://github.com/intel-analytics/ipex-llm
[2] https://github.com/intel/intel-extension-for-pytorch
[3] https://www.intel.com/content/www/us/en/developer/tools/onea...
I'm guessing their marketing department isn't known as the "A-team".
And Nvidia doesn't want to cannibalize its high end chips but putting more memory into consumer ones.
And even if you spend a lot of die space on memory controllers, you can only fit so many GDDR chips around the GPU core while maintaining signal integrity. HBM sidesteps that issue but it's still too expensive for anything but the highest end accelerators, and the ordinary LPDDR that Apple uses is lacking in bandwidth compared to GDDR, so they have to compensate with ginormous amounts of IO silicon. The M4 Ultra is expected to have similar bandwidth to a 4090 but the former will need a 1024bit bus to get there while the latter is only 384bit.
If we're being reasonable and say that you're not using a modern HEDT CPU that costs a couple thousand, the best a consumer botherboard can get right now would be 2x 8x PCIe gen 5 at 32GB/s and one chipset x8 PCIe gen 4 at 16GB/s. I'm not sure if a motherboard like that actually exists but Intel's chipset should allow it; AMD only does x4 to chipset so the third slot is limited by that
If I need 300GB/s memory bandwidth for my workload, that can be accomplished with:
* One RAM chip with 300GB/s
* Two RAM chips with 150GB/s each
* Four RAM chips with 75GB/s each
Etc.
Stepping up from 16GB to 196GB, the bandwidth requirements for each chip go down 10-fold, and you can use much cheaper RAM as a result. And all the signalling requirements relax too.
Much of this discussion presumes a 200GB card would individually need the same capacity to each RAM chip as a 12GB card. This is just false. An A770 or 4060-grade card couldn't keep up with that much data. And if I'm using a small model, I can get the same bandwidth by properly distributing it among RAM chips (which most hardware does automatically).
An A770 or 4060-grade card, with the same total memory capacity as we have today, but 200GB RAM, would allow us to run high-quality LLMs locally or do high-resolution renders. That wouldn't have the same performance as a $200k card, obviously, but for many inferences uses, that's just not very important.
If I were buying for my own uses, I'd want 12x 32GB PC3200 DIMMs for a total of 384GB RAM at $600 for the RAM (say $2k total), with an individual throughput of 25GB/sec and a total throughput of 300GB/sec. I'd be okay with 4060-grade performance. My own uses are a bit niche, and I think for most other people's uses, something with a little more throughput and a little less capacity (48-196GB) might make more sense. But you definitely don't need the same throughput as existing GPU RAM.
If this hypothetical 128GB LLM accelerator was also a capable GPU that would be more interesting but Intel hasn't proven an ability to execute on that level yet.
They already replied with an answer.
https://www.qualcomm.com/news/onq/2023/11/introducing-qualco...
It would be interesting if those saying that a regular GPU with 128GB of VRAM cannot be made would explain how Qualcomm was able to make this card. It is not a big stretch to imagine a GPU with the same memory configuration. Note that Qualcomm did not use HBM for this.
CPUs and GPUs access memory very differently.
Why would it need to be introduced at Apple's high-margin pricing?
Intel doesn't currently have nodes competitive with TSMC or excess capacity in their better processes.
They don't even have competitive capacity for all their CPU needs. They have negative spare capacity overall.
Intel foundry screwed up so badly that Nokia's server division was almost shut down because of Intel Foundry's failure. (imagine being so bad at your job, that your clients go out of business) If Intel client side chose to use Foundry, there just wouldn't be any chips to sell.
Intel literally outsourced their Arrow Lake manufacturing to TSMC because they couldn't fabricate the parts themselves - their 20A (2nm) process node never reached a production-ready state, and was eventually cancelled about a month ago.
Like the brand new Mini that cost 600 USD and went to 500 during Black week.
The good one which is still slower than m4 max is 2200.
If you want the max you need at least a macbook pro starting at 3200 and if you want the better one with 128G RAM it starts at about 5k
Better than most of the pc's out there.
The average desktop isn't exactly a gaming machine. But a corporate box or a low end home desktop with a Core i5 and using the iGPU.
Everyone else wants configurable RAM that scales both down (to 16GB) and up (to 2TB), to cover smaller laptops and bigger servers.
GPUs with soldered on RAM has 500GB/sec bandwidths, far in excess of Apples chips. So the 8GB or 16GB offered by NVidia or AMD is just far superior at vid o game graphics (where textures are the priority)
Apple is doing 800GB/sec on the M2 Ultra and should reach about 1TB/sec with the M4 Ultra, but that's still lagging behind GPUs. The 4090 was already at the 1TB/sec mark two years ago, the 5090 is supposedly aiming for 1.5TB/sec, and the H200 is doing 5TB/sec.
It's pretty expensive though.
The 500GB/sec number is for a more ordinary GPU like the B580 Battlemage in the $250ish price range. Obviously the $2000ish 4090 will be better, but I don't expect the typical consumer to be using those.
There's pretty much a direct opposite scaling between flexibility and performance - dimms > soldered ram > on-package ram > die-interconnects.
We'd be happy to pay 5000 for 128gb from Intel.
Yes, yes, it's not trivial to have a GPU with 128gb of memory with cache tags and so on, but is that really in the same universe of complexity of taking on Nvidia and their CUDA / AI moat any other way? Did Intel ever give the impression they don't know how to design a cache? There really has to be a GOOD reason for this, otherwise everyone involved with this launch is just plain stupid or getting paid off to not pursue this.
Saying all this with infinite love and 100% commercial support of OpenCL since version 1.0, a great enjoyer of A770 with 16GB of memory, I live to laugh in the face of people who claimed for over 10 years that OpenCL is deprecated on MacOS (which I cannot stand and will never use, yet the hardware it runs on...) and still routinely crushes powerful desktop GPUs, in reality and practice today.
You can even find some attractive deals on motherboard/ram/cpu bundles built around grey market engineering sample CPUs on aliexpress with good reports about usability under Linux.
Building a whole new system like this is not exactly as simple as just plugging a GPU into an existing system, but you also benefit from upgradeability of the memory, and not having to use anything like CUDA. llamafile, as an example, really benefits from AVX-512 available in recent CPUs. LLMs are memory bandwidth bound, so it doesn't take many CPU cores to keep the memory bus full.
Another benefit is that you can get a large amount of usable high bandwidth memory with a relatively low total system power usage. Some of AMD's parts with 12 channel memory can fit in a 200W system power budget. Less than a single high end GPU.
I'm a rendering guy, not an AI guy, so I really just want the teraflops, but all GPU users urgently need a 3rd market player.
Believe me, as a lifelong hobbyist-HPC kind of person, I am absolutely dying for such a HBM/fp64 deal again.
https://www.aliexpress.us/item/3256807766813460.html
Doesn't seem like 20x to me. I'm sure spending more than 30 seconds searching could find even better deals.
This is exactly what I meant about Intel's recent launch. Imagine if they went full ALU-heavy on latest TSMC process and packaged 128GB with it, for like, 2-3k Eur. Nvidia would be whipping their lawyers to try to do something about that, not just their engineers.
https://github.com/ryao/llama3.c
My experience is that input processing (prompt processing) is compute bottlenecked in GEMM. AVX-512 would help there, although my CPU’s Zen 3 cores do not support it and the memory bandwidth does not matter very much. For output generation (token generation), memory bandwidth is a bottleneck and AVX-512 would not help at all.
Whenever I see DDR5 memory channels discussed, I am never sure if the speaker is accounting for the 2x 32-bit channels per DIMM or not.
https://www.intel.com/content/www/us/en/products/details/pro...
https://www.qualcomm.com/news/onq/2023/11/introducing-qualco...
They did not put it into the PC parts supply chain for reasons known only to them. That said, it would be awesome if Intel made high memory variants of their Arc graphics cards for sale through the PC parts supply chains.
Sure that was only gddr5 and not gddr6 or lpddr5, but I would have bet we'd be up to 512bit again 10 years down the line..
(I mean supposedly hbm3 has done 1024-2048bit busses but that seems more research or super high end cards, not consumer)
Does M4 Max have 64-byte cache lines?
If they can fetch or flush an entire cache line in a single memory-bus transaction, I wonder if that opens up any additional hardware / performance optimizations.
on the CPU side: 64 bytes at L1, 128 byte cachelines at L2
In the pedantic sense of just literally slapping more on existing boards? No, they might have one empty spot for an extra BGA VRAM chip, but not enough for the gain's we're talking about. But this is absolutely possible, trivially so for someone like Intel/AMD/NVidia, that has full control over the architectural and design process. Is it a switch they flip at the factory 3 days before shipping? No, obviously not. But if they intended this to be the case ~2 years ago when this was just a product on the drawing board? Absolutely. There is 0 technical/hardware/manufacturing reason they couldn't do this. And considering the "entry level" competitor product is the M4 Max which starts at at least $3,000 (for a 128GB equipped one), the margin on pricing more than exists to cover a few hundred extra in ram and extra overhead in higher-layer more populated PCB's.
The real impediment is what you landed on at the end there combined with the greater ecosystem not having support for it. Intel could drop a card that is, by all rights, far better performing hardware than a competing Nvidia GPU, but Nvidia's dominance in API's, CUDA, Networking, Fabric-switches (NVLink, mellanox, bluefield), etc etc for that past 10+ years and all of the skilled labor that is familiar with it would largely render a 128GB Arc GPU a dud on delivery, even if it was priced as a steal. Same thing happened with the Radeon VII. Killer compute card that no one used because while the card itself was phenomenal, the rest of the ecosystem just wasn't there.
Now, if intel committed to that card, and poured their considerable resources into that ecosystem, and continued to iterate on that card/family, then now we're talking, but yeah, you can't just 10X VRAM on a card that's currently a non-player in the GPGPU market and expect anyone in the industry to really give a damn. Raise an eyebrow or make a note to check back in a year? Sure. But raise the issue to get a greenlight on the corpo credit line? Fat chance.
There's a huge number of people in that community that would love to have such a card. How many are actually willing and able to pony up >=$3k per unit? How many units would they buy? Given all of the other considerations that go into making such cards useful and easy to use (as described), the answer is - in intel's mind - nowhere near enough, especially when the financial side of the company's jimmies are so rustled that they sacked Pat G without a proper replacement and nominated some finance bros in as interim CEO's. Intel is ALREADY taking a big risk and financial burden trying to get into this space in the first place, and they're already struggling, so the prospect of betting the house like that just isn't going to fly for the finance bros that can't see passed the next 2 quarters.
To be clear, I personally think there is huge potential value in trying to support the OSS community to, in essence, "crowd source" and speedrun some of that ecosystem by supplying (Compared to the competition) "cheap" cards that aschew the artificial segmentation everyone else is doing and investing in that community. But I'm not running Intel, so while that'd be nice, it's not really relevant.
In Q1 2023, Intel sold 250,000 ARC cards. Sales then collapsed the next quarter. I would expect sales to easily exceed that and be maintained. The demand for high memory GPUs is far higher than many realize. You have professional inferencing operations such as the ones listed at openrouter.ai that would gobble up 128GB VRAM ARC cards for running smaller high context models, much like how you have businesses gobbling up the Raspberry Pi for low end tasks, without even considering the local inference community.
Why not? It doesn't have to be balanced. RAM is cheap. You would get an affordable card that can hold a large model and still do inference e.g. 4x faster than a CPU. The 128GB card doesn't have to do inference on a 128GB model as fast as a 16GB card does on a 16GB model, it can be slower than that and still faster than any cost-competitive alternative at that size.
The extra RAM also lets you do things like load a sparse mixture of experts model entirely into the GPU, which will perform well even on lower end GPUs with less bandwidth because you don't have to stream the whole model for each token, but you do need enough RAM for the whole model because you don't know ahead of time which parts you'll need.
A more sensible alternative would be going with HBM, except good luck getting any capacity for that since it's all being used for the extremely high margin data center GPUs. HBM is also extremely expensive both in terms of the cost of buying the chips and due to it's advanced packaging requirements.
This assumes you use 32Gbit chips, which will likely be available in the near future. Interestingly, the GDDR7 specification allows for 64Gbit chips:
> the GDDR7 standard officially adds support for 64Gbit DRAM devices, twice the 32Gbit max capacity of GDDR6/GDDR6X
https://www.anandtech.com/show/21287/jedec-publishes-gddr7-s...
And if you want to have some real fun, cause "registered GDDR" to be a thing.
Splitting the model up between several GPUs would add a third much worse bottleneck – memory bandwidth between the GPUs. No matter how well you connect them, it'll be slower than transfer within a single GPU.
Still, the fact that you can fit an 8× larger GPU might be worth it to you. It's a trade-off that's almost universally made while training LLMs (sometimes even with the model split down both its width and length), but is much less attractive for inference.
What if you allowed the system to only have a shared memory between every neighboring pair of GPUs?
Would that make sense for an LLM?
Not too many hardware enthusiast site editors have that academic background.
And while fervor can sometimes substitute for education... probably not in microprocessor / system design.
As for articles, IMO, Chips and Cheese is the closest thing we have to Anandtech or RWT in their peak.
https://www.qualcomm.com/news/onq/2023/11/introducing-qualco...
Would someone with “basic ‘knowledge’ of hardware” explain why a GPU cannot be made with the same memory configuration?
Re: Tomasulo's algorithm the other day: https://news.ycombinator.com/item?id=42231284
Cerebras WSE-3 has 44 GB of on-chip SRAM per chip and it's faster than HBM. https://news.ycombinator.com/item?id=41702789#41706409
Intel has HBM2e off-chip RAM in Xeon CPU Max series and GPU Max;
What is the difference between DDR, HBM, and Cerebras' 44GB of on-chip SRAM?
Tomasulo's algorithm also centralizes on a common data bus (the CPU-RAM data bus) which is a bottleneck that must scale with the amount of RAM.
Can in-RAM computation solve for error correction without redundant computation and consensus algorithms?
Can on-chip SRAM be built at lower cost?
Von Neumann architecture: https://en.wikipedia.org/wiki/Von_Neumann_architecture#Von_N... :
> The term "von Neumann architecture" has evolved to refer to any stored-program computer in which an instruction fetch and a data operation cannot occur at the same time (since they share a common bus). This is referred to as the von Neumann bottleneck, which often limits the performance of the corresponding system. [4]
> The von Neumann architecture is simpler than the Harvard architecture (which has one dedicated set of address and data buses for reading and writing to memory and another set of address and data buses to fetch instructions).
Modified Harvard architecture > Comparisons: https://en.wikipedia.org/wiki/Modified_Harvard_architecture
C-RAM: Computational RAM > DRAM-based PIM Taxonomy, See also: https://en.wikipedia.org/wiki/Computational_RAM
SRAM: Static random-access memory https://en.wikipedia.org/wiki/Static_random-access_memory :
> Typically, SRAM is used for the cache and internal registers of a CPU while DRAM is used for a computer's main memory.
/? TOPS/W and FLOPS/W: https://www.google.com/search?q=TOPS%2FW+and+FLOPS%2FW :
- "Why TOPS/W is a bad unit to benchmark next-gen AI chips" (2020) https://medium.com/@aron.kirschen/why-tops-w-is-a-bad-unit-t... :
> The simplest method therefore would be to use TOPS/W for digital approaches in future, but to use TOPS-B/W for analogue in-memory computing approaches!
> TOPS-8/W
> [ IEEE should spec this benchmark metric ]
- "A guide to AI TOPS and NPU performance metrics" (2024) https://www.qualcomm.com/news/onq/2024/04/a-guide-to-ai-tops... :
> TOPS = 2 × MAC unit count × Frequency / 1 trillion
- "Looking Beyond TOPS/W: How To Really Compare NPU Performance" (2023) https://semiengineering.com/looking-beyond-tops-w-how-to-rea... :
> TOPS = MACs * Frequency * 2
> [ { Frequency, NNs employed, Precision, Sparsity and Pruning, Process node, Memory and Power Consumption, utilization} for more representative variants of TOPS/W metric ]
Achieving 128GB VRAM with a 256-bit bus (which seems like a reasonable bus width) would mean some multiple of 8 chips. If Micron, Samsung or SK Hynix made 128Gb GDDR7 chips, then 8 would suffice. The best right now seems 24Gb, although 32Gb seems likely to follow (and it would likely come sooner if a large customer such as Intel asked for it), so they would just need to have 32 chips in a quad rank configuration to achieve 128GB.
This assumes that there is no limit in the GDDR7 specification that prevents quad rank configurations. If there is and it still supports dual rank like GDDR6X did, then a 512-bit bus could be done. It would likely be extremely pricy and require a new chip tape out that has much more IO logic transistors to handle the additional bus width (and IO logic transistor scaling is dead, so the die area would be huge), but it is hypothetically possible. Given how much people are willing to pay for more VRAM, it could make business sense to do.
Even if there is no limit in the GDDR7 specification that prevents quadrank, their memory IO logic would need to support it and if it does not, they would need to redesign that and do a new chip tape out in addition to a new board design. This would also be very expensive, although not as expensive as going to a 512-bit memory interface.
In summary, adding more memory would cost more to do and it would not improve competitiveness in the target market for these cards, which I imagine is the main reason that they do not do it.
By the way, the reason that Nvidia implemented support for 2 chips per channel is because they wanted to be able to reach 48GB VRAM on the workstation variant of the 3090 that is known as the RTX A6000 (non-Ada). I do not know why they used 24x 8Gb chips rather than 12x 16Gb on the 3090, although if I had to guess, it had something to do with rank interleaving.
I've not seen any proposals for buffering LPDDR or GDDR, so an analog to LRDIMMs is not a readily-available technology.
GDDR is the memory technology that operates at the edge of what's possible for per-pin bandwidth. Loading that memory bus down with many ranks is not something we can expect to be achievable by just putting down more pads on the PCB.
That said, I am not an electrical engineer (although I work alongside one and have had a minor role in picking low end components for custom PCBs), I think if Intel were to make a GPU with 128GB VRAM using GDDR7 in the next year or two, the engineer who does the trace routing to make it possible should make a go fund me page for people to send beer money.
In terms of what would have been feasible for Intel to bring to market in 2024, the cheapest option for 128GB capacity would probably have been ~8.5Gb/s LPDDR5x on a 256-bit bus, but to at least match the bandwidth of the chip they just launched, it would have made more sense to use a 512-bit bus and bump the die size back up to ~half the reticle limit like their previous generation die with a 256-bit bus. So they would have had a quite slow but high-capacity GPU with a manufacturing cost equal to at least an RTX 4080, before adding in the cost of all that DRAM. And if they had started working on that chip as soon as LLaMA went public, they might have been able to deliver it by now.
It's no surprise at all that such a risky niche product did not emerge from a division of Intel that is lucky to not have been liquidated yet.
Intel is rumored to have a B770 GPU in development, but it was running late and then was delayed to next year since it had yet to tape out, so they are launching their B580 and B570 graphics cards, which had been ready to go for a while, now. That is why the bus size appears to have dropped across generations. Presumably, if they made a 512-bit bus version, it would be a 9 series card. They certainly left room for it in their lineup, but as far as leaks have been concerned, there is not much hope for one. I do not expect them to use anything other than GDDR7 on their battlemage cards.
As for a high memory ARC card, I am of the opinion that such a product would sell well among the local llama community. There might even be more sales of a high memory ARC card for inference than of the regular ARC cards for gaming given that their discrete graphics sales peaked at 250,000 in Q1 2023 before collapsing, which can be confirmed using the data here:
https://www.tomshardware.com/pc-components/gpus/discrete-gpu...
The market for high memory GPUs is surely bigger than that. That said, Intel is likely pricing their ARC GPUs at a loss after R&D costs are considered. This is likely intended to help them break into a new market, although it has not been going well for them so far. I would guess that they are at least a generation away from profitability.
Intel intends for its Gaudi 3 accelerators to be used for this rather than the ARC line. Those coincidentally have 128GB of RAM, but they use HBM rather than a DDR variant. Qualcomm on the other hand made its own accelerator with 128GB of LPDDR4x RAM:
https://www.qualcomm.com/news/onq/2023/11/introducing-qualco...
If my math is right, Qualcomm went with a 1024-bit memory bus and some incorrect rounding (rounding 137.5 to 138 before multiplying by 4) to reach their stated bandwidth figure. Qualcomm is not selling it through the PC parts supply chain, so I have no clue how much it costs, but I assume that it is expensive. I assume that they used LPDDR4x to be able to build a product since they were too late in securing HBM supply and even if they did, they would not be able to scale production to meet demand growth since Nvidia is buying all of the HBM that it can.
That is objectively false. See, for instance, V-color’s threadripper RAM[0]. If 96GB quad-rank modules @ 6000Mhz in octo-channel counts as “barely operating” maybe we have different definitions of operation requirements.
As a side note, their quad-channel 3-rank RAM [1] hits 8000MHz, out of the box. Admittedly only 24GB modules, but still.
[0] https://v-color.net/products/ddr5-ocrdimm-amd-wrx90-workstat... [1] https://v-color.net/products/ddr5-oc-rdimm-amd-trx50-worksta...
I think the high memory local inference stuff is going to come from "AI enabled" cpus that share the memory in your computer. Apple is doing this now, but cheaper options are on the way. As a shape its just suboptimal for graphics, so it doesn't make sense for any of the gpu vendors to do it.
Unfortunately I don't think either Intel or AMD makes a CPU that supports quad channel RAM at a decent price.
My 7900 XTX does 120TFlops.
To match that, you would need to scale that CPU up to either 2048 cores, 2KB per register (still one-cycle!) or 64Ghz.
I guess if you had 1024-bit registers and 8Ghz, you could get away with only 240 cores. Good luck thermal dissipating that btw. To reverse an opinion I'm seeing in this thread, at that point your CPU starts looking more like a GPU by necessity.
For inference, prompt processing is compute intensive, while token generation is memory bandwidth bound. The differences in memory bandwidth between CPUs and GPUs tend to be more profound than the difference in compute.
https://overclockers.ru/st/legacy/blog/428111/424644_O.jpg
I have been working on doing inference on a Ryzen 7 5800X lately and I have had good results:
https://github.com/ryao/llama3.c/blob/master/run.c
Running on a GPU like my 3090 Ti will likely outperform it by two orders of magnitude, but I have managed to push the needle slightly on the state of the art performance for prompt processing on my CPU. I suspect an additional 15% improvement is possible, but I do not expect to be able to realize it. In any case, it is an active R&D project that I am doing to learn how these things work.
Finally to answer your question, I have no good answers for you (or more specifically, answers that I like). I have been trying to think of ways to do fast local inference on high end models cost effectively for months. So far, I have nothing to show for it aside from my R&D into CPU llama 3 inference since none of my ideas are able to bring hardware costs needed for llama 3.1 405B below $10,000 with performance at an acceptable level. My idea of an acceptable performance level is 10 tokens per second for token generation and 4000 tokens per second for prompt processing, although perhaps lower prompt processing performance is acceptable with prompt caching.
Once you have offloaded flash attention, you're back to GEMV having a memory bottleneck. GEMV does a single multiplication and addition per parameter. You can add as many EXAFLOPs as you want, it won't get faster than your memory.
And that 500GB/sec is pretty low for a gpu, its like a 4070 but the memory alone would add $500+ to the cost of the inputs, not even counting the advanced packaging (getting those bandwidths out of lpddr needs organic substrate).
It's not that you can't, just when you start doing this it stops being like a graphics card and becomes like a cpu.
https://www.tomshardware.com/pc-components/dram/samsung-outl...
I am not sure why they can already do stacking for HBM, but not GDDR and DDR. My guess is that it is cost related. I have heard that HBM reportedly costs 3 times more than DDR. Whatever they are doing to stack it now that is likely much more expensive than their planned 3D fabrication node.
As far as I understand it, it gives you 64 GiB of HBM per socket.
But obvoiusly they don't.
And for reasons: NVidia has worked on CUDA for ages, do you believe they just replace this whole thing in no time?
You can't just "add more ram" to GPUs and have them work the same way. Memory access is completely different than on CPUs.
llama.cpp is just inference, not training, and the CUDA backend is still the fastest one by far. No one is even close to matching CUDA on either training or inference. The closest is AMD with ROCm, but there's likely a decade of work to be done to be competitive.
You're talking about completely different things here.
It's fine if you're doing a few requests at home, but if you're actually serving AI models, CUDA is the only reasonable choice other than ASICs.
A dev, if they want to run local models, probably run something which just fits on a proper GPU. For everything else, everyone uses an API key from whatever because its fundamentaly faster.
IF a affordable intel GPU would be relevant faster for inferencing, is not clear at all.
A 4090 is at least double the speed of Apples GPU.
I have a 4090, PCIe 3x16, DDR4 RAM.
oobabooga/text-generation-webui
using exllama
I can load 30B 4bit GPTQ models and use full 2048 context
I get 30-40 tokens/s
[1] https://old.reddit.com/r/LocalLLaMA/comments/14gdsxe/optimal...Keep NVIDIA for training and Intel/AMD/Cerebras/… for interference.
And it needs liquid cooling.
You don't just plugin intel cards 'out of the box'.
Inference is also a much smaller market right now, but will likely be overtaken later as we have more people using the models than competing to train the best one.
Would anyone choose llama.cpp's training tools to do serious work? No. Do they exist and work, yes.
Or you can just buy Nvidia.
In this case, the die I/O limits precludes more than a reasonable number of DDR channels.
High speed, low latency server grade DDR5 is around $800-$1600 for 128GB. Triple that for $2400 - $4800 just for the memory. Still need the GPUs/APUs, card, VRMs, etc.
Even the nVidia H100 with "only" 94GB starts at $30k...
Their last quarter was $35b in sales and $26b in gross profit ($21.8b op income; 62% op income margin vs sales).
Visa is notorious for their extreme margin (66% op income margin vs sales) due to being basically a brand + transaction network. So the fact that a hardware manufacturer is hitting those levels is truly remarkable.
It's very clear that either AMD or Intel could accept far lower margins to go after them. And indeed that's exactly what will be required for any serious attempt to cut into their monopoly position.
You misunderstand why and how Nvidia is a monopoly. Many companies make GPUs, and all those GPUs can be used for computation if you develop compute shaders for them. This part is not the problem, you can already go buy cheaper hardware that outperforms Nvidia if price is your only concern.
Software is the issue. That's it - it's CUDA and nothing else. You cannot assail Nvidia's position, and moreover their hardware's value, without a really solid reason for datacenters to own them. Datacenters do not want to own GPUs because once the AI bubble pops they'll be bagholders for Intel and AMD's depreciated software. Nvidia hardware can at least crypto mine, or be leased out to industrial customers that have their own remote CUDA applications. The demand for generic GPU compute is basically nonexistent, the reason this market exists at all is because CUDA exists, and you cannot turn over Nvidia's foothold without accepting that fact.
The only way the entire industry can fuck over Nvidia is if they choose to invest in a complete CUDA replacement like OpenCL. That is the only way that Nvidia's value can be actually deposed without any path of recourse for their business, and it will never happen because every single one of Nvidia's competitors hate each other's guts and would rather watch each other die in gladiatorial combat than help each other fight the monster. And Jensen Huang probably revels in it, CUDA is a hedged bet against the industry ever working together for common good.
It's impossible to assail their monopoly without utilizing far lower prices, coming up under their extreme margin products. It's how it is almost always done competitively in tech (see: ARM, or Office (dramatically undercut Lotus with a cheaper inferior product), or Linux, or Huawei, or Chromebooks, or Internet Explorer, or just about anything).
Note: I never said lower prices is all you'd need. Who would think that? The implication is that I'm ignorant of the entire history of tech, it's a poor approach to discussion with another person on HN frankly.
In the case of ARM, Office, Linux, Huawei, and ChromeOS, these were all actual alternatives to the incumbent tools people were familiar with. You can directly compare Office and Lotus because they are fundamentally similar products - ARM had a real chance against x86 because wasn't a complex ISA to unseat. Nvidia is not analogous to these businesses because they occupy a league of their own as the provider of CUDA. It's not exaggeration to say that they have completely seceded from the market of GPUs and can sustain themselves on demand from crypto miners and AI pundits alone.
AMD, Intel and even Apple have bigger things to worry about than hitting an arbitrary price point, if they want Nvidia in their crosshairs. All of them have already solved the "sell consumer tech at attractive prices" problem but not the "make it complex, standardize it and scale it up" problem.
So many billions of dollars and no one is even 1% close to displacing CUDA in any meaningful way. ZULDA is dead. ROCM is a meme, Scale is a meme. Either you use CUDA or you don't do meaningful AI work.
I don't personally think CUDA is impossible to replace - but I do think that everyone capable of replacing CUDA has been ignoring it recently. Nvidia's role as the GPGPU compute people is secure for the foreseeable future. Apple wants to design simpler GPUs, AMD wants to design cheaper GPUs, and Intel wants to pretend like they can compete with AMD. Every stakeholder with the capacity to turn this ship around is pretending like Nvidia doesn't exist and whistling until they go away.
Intel seems to have thrown their weight behind SYCL, which is an open standard intended to compete with CUDA. Its not clear there has been much interest from other hardware vendors though.
They processed $12T in payments last year (almost a billion payments per day), with a net revenue of $32B. That's a gross transaction margin of 0.26% and their GAAP net income was half that, about 0.14%. [1]
They're just a transaction network, unlike say Amex which is both an issuer and a network. Being just the network is more operationally efficient.
It's quite simple. Divide revenue minus costs by revenue. Transaction volume isn't revenue. Visa only gets the transaction fee.
Even if I give you the benefit of the doubt and do a proper interpretation of the number you've arrived at, its meaning is quite different and quite off topic from this discussion. What you have calculated is the total share of costs that Visa represents in that 12 trillion dollar part of the economy. It is like saying Visa's share of GDP is 0.1%.
As my mother used to say if you have nothing nice to say you're better off staying quiet ;)
also i didn't know there was a 192GB amd GPU.
For local inference, 7900 XTX has 24 GB of VRAM for less than $1000.
At what threshold of VRAM would you start being interested in MI?
Models are getting larger, not smaller. This is why H200 has more memory, but the same exact compute. MI300x vs. MI325x... more memory, same compute.
This does not consider that the board of directors would crucify Lisa Su if she authorized the use of HBM on a consumer product while it is supply constrained and there is enterprise demand for products using it. AMD can only get a limited amount of it and what they do get is not enough for enterprise demand where AMD has extremely healthy margins.
Even if they by some miracle turned a profit on a $2000 consumer card with 192GB HBM, every sale would have a massive opportunity cost and effectively would be a loss in the eyes of the board of directors.
Meanwhile, Nvidia would be unaffected because AMD could not produce very many of these.
If Intel or AMD sold a niche product with 48GB RAM even at a loss, but hit high-end consumer pricing, there would be a flood of people doing various AI work to buy it. The end result would be that parts of NVidia's moat would start draining rather quickly, and AMD / Intel would be in a stronger position for AI products.
I use NVidia because when I bought AMD during the GPU shortage, ROCm simply didn't work for AI. This was a few years back, but I was burned badly enough that I'm unlikely to risk AMD again for a long, long time. Unused code sits broken, and no ecosystem gets built up. A few years later, things are gradually improving for AMD for the kinds of things I wanted to do years ago, but all my code is already built around NVidia, and all my computers have NVidia cards. It's a project with users, and all those users are buying NVidia as well (even if just for surface dependencies, like dev-ops scripts which install CUDA). That, times thousands of projects, is part of NVidia's moat.
If I could build a cheap system with around 200GB, that would be incentive for me to move the relatively surface dependencies to work on a different platform. I can buy a motherboard with four PCI slots, and plug in four 48GB cards to get there. I'd build things around Intel or AMD instead.
The alternative is NVidia would start shipping competitive cards. If they did that, their high-end profit margins would dissolve.
The breakpoints for inference functionality are really at around 16GB, 48GB, and 200GB, for various historical reasons.
Even if AMD dropped the price to $2000, you could not be able to build a system with one of these cards. You cannot buy these cards at their current 5 digit pricing. The idea that you could buy it if they dropped the price to $2000 is a fantasy, since others would purchase the supply long before you have a chance to purchase one, just like they do now.
AMD is already selling out at the current 5 digit pricing and Nvidia is not affected, since Nvidia is selling millions of cards per year and still cannot meet demand while AMD is selling around 100,000. AMD dropping the price to $2000 would not harm Nvidia in the slightest. It would harm AMD by turning a significant money maker into a loss leader. It would also likely result in Lisa Su being fired.
By the way, the CUDA moat is overrated since people already implement support for alternatives. llama.cpp for example supports at least 3. PyTorch supports alternatives too. None of this harms Nvidia unless Nvidia stops innovating and that is unlikely to happen. A price drop to $2000 would not change this.
HGX B200: 36 petaflops at FP16. 14.4 terabytes/second bandwidth.
RX4060 (similar to Intel): 15 teraflops at FP16. 272 gigabytes/second bandwidth
Hmmm.... Note the prefixes (peta versus tera)
A lot of that is apples-to-oranges, but that's kind of the point. It's a different market.
A low-performance high-RAM product would not cut into the server market since performance matters. What it would do is open up a world of diverse research, development, and consumer applications.
Critically, even if it were, Intel doesn't play in this market. If what you're saying were to happen, it should be a no-brainer for Intel to launch a low-cost alternative. That said, it wouldn't happen. What would happen is a lot of business, researchers, and individuals would be able to use ≈200GB models on their own hardware for low-scale use.
> By the way, the CUDA moat is overrated since people already implement support for alternatives.
No. It's not. It's perhaps overrated if you're building a custom solution and making the next OpenAI or Anthropic. It's very much not overrated if you're doing general-purpose work and want things to just work.
https://www.nvidia.com/en-us/data-center/hgx/ https://www.techpowerup.com/gpu-specs/geforce-rtx-4060.c4107
The deeper problem is that the market for this is probably incredibly niche.
It is why you can only see GPUs with 24GB of memory at the moment.
HBM2 can handle 64GB ( 4 x 8GB Stack ) ( Total capacity 128GB )
HBM3 can handle 192GB ( 4 x 24GB Stack ) ( Total capacity 384GB )
You can not do this with GDDR6.
The problem you experience with GDDR6 and channels is width. You get higher in size but you'll never be able to fill the memory fast enough to be cost effective. Also why HBM exists. Lets say the top is 960GB/s for an A6000 for GPU memory. HBM3 is 3.35 TB/s.
If Intel wanted to make better GPUs it needs to switch to HBM.
You say nope and then proceed to say exactly what I said. When if the ICs are limited to 16Gbit density for now, the 24GB limit you mentioned was wrong, since you assumed single rank was the highest possible.
As for HBM, that is earmarked for their gaudi 3 line. There is no chance of it being put into a consumer product, as it would be like selling gold at pyrite pricing.
Because, no. It's not as simple as that.
NVIDIA has a complete ecosystem now. They have cards. They have cards of cards (platforms), which they produce, validate and sell. They have NVLink crossbars and switches which connects these cards on their card of cards with very high speeds and low latency.
For inter-server communication they have libraries which coordinate cards, workloads and computations.
They bought Mellanox, but that can be used by anyone, so there's no lock-in for now.
As a tangent, NVIDIA has a whole set of standards for pumping tremendous amount of data in and out of these mesh of cards. Let it be GPU-Direct storage or specialized daemons which handle data transfers on and off cards.
If you think that you can connect n cards on PCIe bus and just send workloads to them and solve problems magically, you'll hurt yourself a lot, both performance and psychology wise.
You have to build a stack which can perform these things with maximum possible performance to be able to compute with NVIDIA. It's not just emulating CUDA, now. Esp., on the high end of the AI spectrum (GenAI, MultiCard, MultiSystem, etc.).
For other lower end, multi-tenant scenarios, they have card virtualization, MIG, etc. for card sharing. You have to complete on that, too, for cloud and smaller applications.
It's like the most profitable set of products in tech. You have companies like Meta, MSFT, Amazon, Google etc spending $5B every few years buying this hardware.
Edit: I can’t sort this out. Where did all the money go?
However, AMD is coming for them because a couple of high profile supercomputer centers (LUMI, Livermore, etc.) are using Instinct cards and pouring money to AMD to improve their cards and stack.
I have not used their (Instinct) cards, yet, but their Linux driver architecture is way better than NVIDIA.
I wonder where you got your information on AMD’s “Linux driver architecture”. It is reportedly a mess:
https://news.ycombinator.com/item?id=34832660
So far, I have been very happy with Nvidia’s Linux drivers.
If you're doing inference on a server; MIG comes into play. If you're doing inference on a larger cloud, GPU-direct storage comes into play.
It's all modular.
GPU-Direct is about pumping data from storage devices to cards, esp. from high speed storage systems across networks.
MIG actually shares a single card to multiple instances, so many processes or VMs can use a single card for smaller tasks.
Nothing I have written in my previous comment is related to inter-card, inter-server communication, but all are related to disk-GPU, CPU-GPU or RAM-CPU communication.
Edit: I mean, it's not OK to talk about downvoting, and downvote as you like but, I install and enable these cards for researchers. I know what I'm installing and what it does. C'mon now. :D
Obviously that is only one piece of software, but its a certainly a useful one if you are using one of the many LLMs it supports.
I had the 16GB arc, and it was able to run inference at the speed i expected, but twice as many per batch as my 8GB card, which i think is about what you'd expect.
once the model is on the card, there's no "disk" anymore, so having more vram to load the model and the tokenizer and whatever else on means there's no disk, and realistically when i am running loads on my 24GB 3090 the CPU is maybe 4% over idle usage. My bottleneck, as it stands, to running large models is vram, not anything else.
If i needed to train (from scratch or whatever) i'd just rent time somewhere, even with a 128GB card locally, because obviously more tensors is better.
and you're getting downvoted because there's literally lm studio and llama.cpp and sd-webui that run just fine for inference on our non-dc, non-nvlink, 1/15th the cost GPUs.
If there's a competing platform that hobbyists can tinker with, the ecosystem can improve quite rapidly, especially when the competing platform is completely closed and hobbyists basically are locked out and have no alternative.
On the contrary. You really don't know how I love and prefer open source and love a more leveling playing field.
> If there's a competing platform that hobbyists can tinker with...
AMD's cards are better from hardware and software architecture standpoint, but the performance is not there yet. Plus, ROCm libraries are not that mature, but they're getting there. Developing high performance, high quality code is deceivingly expensive, because it's very heavy in theory, and you fly very close to the metal. I did that in my Ph.D., so I know what it entails. So it requires more than a couple (hundred) hobbyists to pull off (see the development of Eigen linear algebra library, or any high end math library).
Some big guns are pouring money into AMD to implement good ROCm libraries, and it started paying off (Debian has a ton of ROCm packages now, too). However, you need to be able to pull it off in the datacenter to be able to pull it off on the desktop.
AMD also needs to be able to enable ROCm on desktop properly, so people can start hacking it at home.
> especially when the competing platform is completely closed...
NVIDIA gives a lot of support to universities, researchers and institutions who play with their cards. Big cards may not be free, but know-how, support and first steps are always within reach. Plus, their researchers dogfood their own cards, and write papers with them.
So, as long as papers got published, researchers do their research, and something got invented, many people don't care about how open source the ecosystem is. This upsets me a ton, but when closed source AI companies and researchers who forget to add crucial details to their papers so what they did can't be reproduced don't care about open source, because they think like NVIDIA. "My research, my secrets, my fame, my money".
It's not about sharing. It's about winning, and it's ugly in some aspects.
I'm one of those academics. You've got it all wrong. So many people care about open source. So many people carefully release their code and make everything reproducible.
We desperately just want AMD to open up. They just refuse. There's nothing secret going on and there's no conspiracy. There's just a company that for some inexplicable reason doesn't want to make boatloads of money for free.
AMD is the worst possible situation. They're hostile to us and they refuse to invest to make their stuff work.
Software wise, maybe. But you can't change AMD's hardware with a magic wand, and that's where a lot of CUDA's optimizations come from. AMD's GPU architecture is optimized for raster compute, and it's been that way for decades.
I can assure you that AMD does not have a magic button to press that would make their systems competitive for AI. If that was possible it would have been done years ago, with or without their consent. The problem is deeper and extends to design decisions and disagreement over the complexity of GPU designs. If you compare AMD's cards to Nvidia on "fair ground" (eg. no CUDA, only OpenCL) the GPGPU performance still leans in Nvidia's favor.
That said, for hobbyist inference on large pretrained models, I think there is an interesting set of possibilities here: maybe a number of operations aren't optimized, and it takes 10x as long to load the model into memory... but all that might not matter if AMD were to be the first to market for 128GB+ VRAM cards that are the only things that can run next-generation open-weight models in a desktop environment, particularly those generating video and images. The hobbyists don't need to optimize all the linear algebra operations that researchers need to be able to experiment with when training; they just need to implement the ones used by the open-weight models.
But of course this is all just wishful thinking, because as others have pointed out, any developments in this direction would require a level of foresight that AMD simply hasn't historically shown.
It would be a huge shift for them. To go from preferring some (sometimes not quite reached) metric, to, perhaps rightly play the 'reformed underdog'. Commoditize Big-Memory ML Capable GPUs, even if they aren't quite as competitive as the top players at first.
Will the other players respond? Yes. But ruin their margin. I know that sounds cutthroat[1] but hey I'm trying to hypothetically sell this to whomever is taking the reigns after Pat G.
> NVIDIA gives a lot of support to universities, researchers and institutions who play with their cards. Big cards may not be free, but know-how, support and first steps are always within reach. Plus, their researchers dogfood their own cards, and write papers with them.
Ideally they need to do that too. Ideally they have some 'high powered' prototypes (e.x. lets say they decide a 2-gpu per card design with an interlink is feasible for some reason) to share as well. This may not be be entirely ethical[1] in this example of how a corp could play it out, again it's a thought experiment since intel has NOT announced or hinted at a larger memory card anyway.
> AMD also needs to be able to enable ROCm on desktop properly, so people can start hacking it at home
AMD's driver story has always been a hot mess, My desktop won't behave with both my onboard video and 4060 enabled, every AMD card I've had winds up with some weird firmware quirk one way or another... I guess I'm saying their general level of driver quality doesn't lend to hope they'll fix dev tools that soon...
[0] - https://old.reddit.com/r/LocalLLaMA/comments/12khkka/running...
[1] - As you said, it's about winning and it can get ugly.
Point of the OP is this is entirely possible with even an iGPU if only we have the RAM. nVidia should be irrelevant for local inference.
See the precompute_input_logits() and forward() functions here:
https://github.com/ryao/llama3.c/blob/master/run.c#L520
As a preface, precompute_input_logits() is really just a generalized version of the forward() function that can operate on multiple input tokens at a time to do faster input processing, although it can be used in place of the forward() function for output generation just by passing only a single token at a time.
Also, my apologies for the code being a bit messy. matrix_multiply() and batched_matrix_multiply() are wrappers for GEMM, which I ended up having to use directly anyway when I needed to do strided access. Then matmul() is a wrapper for GEMV, which is really just a special case of GEMM. This is a work in progress personal R&D project that is based on prior work others did (as it spared me from having to do the legwork to implement the less interesting parts of inferencing), so it is not meant to be pretty.
Anyway, my purpose in providing that link is to show what is needed to do inferencing (on llama 3). You have a bunch of matrix weights, plus a lookup table for vectors that represent tokens, in memory. Then your operations are:
* memcpy()
* memset()
* GEMM (GEMV is a special case of GEMM)
* sinf()
* cosf()
* expf()
* sqrtf()
* rmsnorm (see the C function for the definition)
* softmax (see the C function for the definition)
* Addition, subtraction, multiplication and division.
I specify rmsnorm and softmax for completeness, but they can be implemented in terms of the other operations.If you can do those, you can do inferencing. You don’t really need very specialized things. Over 95% of time will be spent in GEMM too.
My next steps likely will be to figure out how to implement fast GEMM kernels on my CPU. While my own SGEMV code outperforms the Intel MKL SGEMV code on my CPU (Ryzen 7 5800X where 1 core can use all memory bandwidth), my initial attempts at implementing SGEMM have not fared quite so well, but I will likely figure it out eventually. After I do, I can try adapting this to FP16 and then memory usage will finally be low enough that I can port it to a GPU with 24GB of VRAM. That would enable me to do what I say is possible rather than just saying it as I do here.
By the way, the llama.cpp project has already figured all of this out and has things running on both GPUs and CPUs using just about every major quantization. I am rolling my own to teach myself how things work. By happy coincidence, I am somehow outperforming llama.cpp in prompt processing on my CPU but sadly, the secrets of how I am doing it are in Intel’s proprietary cblas_sgemm_batch() function. However, since I know it is possible for the hardware to perform like that, I can keep trying ideas for my own implementation until I get something that performs at the same level or better.
Perhaps you can reverse engineer it?
Not really correct for training - training has a lot of all-to-all problems, so hierarchical reduction is useful but doesn't really solve the incast problem - Nvlink _bandwidth_ is less of an issue than perhaps the SHARP functions in the NVLink switch ASICs.
Start by being a "second vendor" for huge customers of NVIDIA that want to foster competition, as well as a few others willing to take risks, and build from there.
Even within that they seem to have difficulty expanding to new areas, remember edison?
I agree, just for my PC, something that'd enable small devs to create interesting foundation model apps that'd deploy to users using this local AI cards to run these new Apps.
I think there have to be a couple of killer apps that run "OK" with CPU or GPU, but would run tremendously better with such a card.
The training can come after, with some inference and runtime optimizations on the software stack.
I'm not sure your argument stands when it comes to OP's idea of a single card with 128GB VRAM. This would be enough to run ~180B models with reasonable quantization --we're not near maxing out the capability of 180B yet (see the latest 32B models performing near public SOTA).
This indeed would push rapid and wide adoption and be quite disruptive. But sure, it wouldn't instantly enable competitive training of 405B models.
Wow what he said is way above your head! Please reread what he wrote.
https://github.com/ryao/llama3.c
Inference workloads are easy to parallelize to N cards with minimal connectivity between them. The Nvlink crossbars and switches just are not needed.
In particular, inference can be divided into two distinct phases, which are input processing (prompt processing) and output generation (token generation). They are remarkably different in their requirements. Input processing is compute bound via GEMM operations while output generation is memory bandwidth bound via GEMV operations. Technically, you can do the input processing via GEMV too by processing 1 token at a time, but that is slow, so you do not want to do that. Anyway, these phases can be further subdivided into the model’s layers. You can have 1 GPU per layer with the logits passing from GPU to GPU in a pipeline. The GPUs just need the layer’s weights and the key-value cache for all of the tokens in that layer in memory to be able to work effectively. For llama 3.1 405B, there are 126 layers, so that is up to 126 GPUs.
That is of course slightly slower than if you just had 1 GPU with an incredible amount of VRAM, but you can always have more than one query in flight to get better than 1 GPU’s worth of performance from this pipeline approach. There are other ways of doing parallelization too, such as having output processing use GEMM to do multiple queries in parallel. This would be what others call batching, although I am only interested in doing 1 query at a time right now, so I have not touched it.
In essence, you can connect n cards on PCIe and have them solve inferencing problems magically, with the right software. Training is a different matter and I cannot comment on it as I have not studied it yet.
From my understanding that's certainly possible to do without the latency hurting much with large batching between inference layers
* the model dimension
* how many bits per variable are used by your quantization
* how many tokens are being processed per step (input processing can do all N input tokens simultaneously and output processing can only do 1 at a time when doing a single query)
* how many times you split the model layers across multiple GPUs
The model dimensions are: * 4096 for llama 3/3.1 8B.
* 8192 for llama 3/3.1 70B.
* 16384 for llama 3.1 405B.
The model layers are: * 32 for llama 3/3.1 8B
* 80 for llama 3/3.1 70B
* 126 for llama 3.1 405B
The amount of data that needs to be transferred for each split is surprisingly small. Each time you move the calculation of a subsequent layer to a different GPU, you need to transfer an array that is of size model_dimension * num_tokens * bits_per_variable. Then this reduces to a classic network transfer time problem, where you consider both time until the first byte arrives and the transfer time until the last byte arrives. Reality will likely be longer than that idealized scenario, especially since you need to send a signal saying to begin computing.Input processing can tackle so many tokens simultaneously that it probably is not worth thinking too much about this penalty there. Output processing is where the penalty is more significant, since you will incur these costs for every token. Let’s say we are doing fp16 or bf16 on llama 3 8B. Then we need to transfer 8KB every time we move the calculation for another layer to another GPU. If you use RDMA and do this over 10GbE, the transfer time would be 6.4 microseconds. If we assume the time to first byte and time to do a signal to begin processing is 3.6 microseconds combined (chosen to round things up), then we get a penalty of 10 microseconds per split, per token. If you are doing 60 tokens per second and split things across 4 GPUs over the network, you have a penalty of 30 microseconds per token. It will run about 0.003% slower and you are not going to notice this at all. Assuming 10GbE with RDMA is somewhat idealized, although I needed to pick something to give some numbers.
In any case, the equation for figuring what factor slower it would be is 1 / (1 + time to do transfers and trigger processing per each token in seconds). That would mean under a less ideal situation where the penalty is 5 milliseconds per token, the calculation will be ~0.99502487562 times what it would have been had it been done in a hypothetical single GPU that has all of the VRAM needed, but otherwise the same specifications. This penalty is also not very noticeable.
In conclusion, you are right. I actually recall seeing a random YouTube video talking about a software project that does clustered interferencing, so people are already doing this. Unfortunately, I do not remember the project name, channel name or video name.
Which therefore means that cards that can only do inference, are fungible. You don’t want to spend CapEx on getting into a new LOB just to sell something fungible.
All the gigantic GPU clusters that you can sell a million at a time to a bigcorp under a high-margin service contract, meanwhile, are training clusters. Nvidia’s market cap right now is fundamentally built on the model-training “space race” going on between the world’s ~15 big AI companies. That’s the non-fungible market.
For Intel to see any benefit (in stock price terms) from an ML-accelerator-card LOB, it’d have to be a card that competes in that space. And that’s a much taller order.
https://www.intel.com/content/www/us/en/products/details/pro...
Coincidentally, it has 128GB of RAM. However, it is not a GPU, is designed to do training too and uses expensive HBM.
Modern GPUs can do more than inference/training and the original poster asked about a GPU with 128GB of RAM, not a card that can only do inferencing as you described. Interestingly, Qualcomm made its own card targeted at only inferencing with 128GB of RAM without using HBM:
https://www.qualcomm.com/news/onq/2023/11/introducing-qualco...
They do not sell it through PC parts channels so I do not know the price, but it is exactly what you described and it has been built. Presumably, a GPU with the same memory configuration would be of interest to the original poster.
Sun (Sparc) and HP (PA-RISC) used to own most of the server market in 1990, but lost most of it to x86 by 2000. Few people had a Sun box with Solaris, but tons of people had access to a PC with Linux, which was inferior in many ways, but well-known and much less locked-up.
Nvidia is starting to sound like a house of cards to me.
Apple is the king right now of local LLM inference, just because of their unified memory architecture meaning that people can get large amounts of "VRAM" (since all RAM is VRAM). They're not as fast as Nvidia — not even close to an H100, for example. But they don't need to be. No consumer can afford an H100, but they can afford a Mac.
I do think one challenge is, AFAIK with most GDDR5/6 there's a density issue that requires either larger memory bus paths or other additional complexity to support large sizes.
That said, the lack of even a 16GB variant is sus.
I'll take some copium in that maybe they're trying to solve the 'size' issue somehow and are just making sure whatever system they use isn't gonna be an i820 MTH debacle before they pull the trigger on announcing it.
https://www.qualcomm.com/news/onq/2023/11/introducing-qualco...
I have no idea how much it costs. They do not sell it via PC parts channels.
Rather than go all in with 128GB, they could test the waters easily with a cheap 32GB offering and take ot from there.
Beyond that, you'd have to move to GDDR7 (which has 24Gb/3GB chips incoming) or to HBM stacks, but at that point you're well beyond a "basic GPU". I think the only way you could get to 128GB would be either using regular (LP)DDR or HBM.
Note, Apple M chips have weak GPUs with decent MBW and large memory capacities (up to 192GB @ 800GB/s for an M2 Ultra, launched mid 2023) and have not been a major CUDA threat so I don't think your hypothesis actually stands up.
besides the clear limitation of the memory technology they are using compared to the nvidia's enterprise solution, for such large GPU chips that could really make use of such memory, they need to make binning possible by selling cut-down versions of them as well.
nvidia can pull this off because they can sell lower-end chips at the same time. intel is barely making a dent in sales, and making a high-end chip will only be very risky, at the cost of potentially benefiting a niche crowd.
> put them as a major CUDA threat. that is a software/ecosystem problem, which hardware alone cannot solve. for all the devs that use Macs, even in AI it is only about inference at the moment. nobody is coming at CUDA for training for the near future. amd tried and failed plenty already.
The lack of quantified stats on the marketing pages tells me Intel is way behind.