Intel's Battlemage Architecture
chipsandcheese.com
chipsandcheese.com
Just a nit: one step up (RX 7600 XT) comes with 16GB memory, although in clamshell configuration. With the B580 falling inbetween the 7600 and 7600 XT in terms of pricing, it seems a bit unfair to only compare it with the former.
- RX 7600 (8GB) ~€300
- RTX 4060 (8GB) ~€310
- Intel B580 (12GB) ~€330
- RX 7600 XT (16GB) ~€350
- RTX 4060 Ti (8GB) ~€420
- RTX 4060 Ti (16GB) ~€580*
*Apparently this card is really rare plus a bad value proposition, so it is hard to find
Not sure where you got that 350 EUR number for B580?
For example:
https://www.mindfactory.de/product_info.php/12GB-ASRock-Inte... (~327 EUR)
https://www.overclockers.co.uk/sparkle-intel-arc-b580-guardi... (~330 EUR)
https://www.inet.se/produkt/5414587/acer-arc-b580-12gb-nitro... (~336 EUR)
Though the market is particularly bad here, because an RTX 3060 12 GB (not Ti) costs between 310 - 370 EUR and an RX 7600 XT is between 370 - 420 EUR.
Either way, I'm happy that these cards exist because Battlemage is noticeably better than Alchemist in my experience (previously had an A580, now it's my current backup instead of the old RX 570 or RX 580) and it's good to have entry/mid level cards.
The reviews will say "decent value at RRP" but Intel cards never ever sell anywhere near RRP meaning that when it comes down to it you're much better off not going Intel.
I feel like reviews should all acknowledge this fact by now. "Would be decent value at RRP but not recommended since Intel cards are always %50 over RRP".
Central Computers in the SF Bay Area keeps them at MSRP. They may not be in stock online, but the stores frequently have stock, especially San Mateo.
Not useful to people outside the area, but then Microcenter also sells at MSRP. So there are non-scalping stores out there.
The trick is to jump on the stock when it arrives.
https://www.newegg.com/p/pl?d=b580&N=8000 to see sold by newegg stock.
If the 4060 was the 3080-for-400USD that everyone actually wants, that'd be a different story. Fortunately, its nonexistence is a major contributor to why the B580 can even be a viable GPU for Intel to produce in the first place.
But the release of those cards was during Covid pricing weirdness times. I scored a 3070 Ti at €650, whilst the 3060 Ti's that I actually wanted were being sold for €700+. Viva la Discord bots.
This card still seems like a bad proposition. It's roughly similar performance to the 11GB 2080 Ti for double the price. You'd have to really want that extra 5GB.
https://www.bestbuy.com/site/pny-nvidia-geforce-rtx-4060-ti-...
Most people who want the 4060 Ti 16GB is because they want the 16GB for running LLMs. So yes, they really want that extra 5GB.
I'm actually tempted, but I don't know if I should go for a Mac Studio M1 Max 64GB for $1350 (ebay) or build a PC around a GPU. I think the Mac makes a lot of sense.
I wonder if the B580 will drop to MSRP at all, or if retailers will just keep it slotted into the greater GPU line-up the way it is now and pocket the extra money.
The RTX 1060 came with 6GB of VRAM. Four generations later, the 5060 comes with only 2GB more.
I suspect NVidia does not want consumer cards to eat into those lucrative data centre profits?
The cost of 1GB of VRAM is $2.30 see https://www.dramexchange.com
Those parts competed with the RX480 with 8GB of memory so NVidia was behind AMD at that price point.
AMD had not been competing with the *80/Ti cards at this point for a few generations and stuck with that strategy through today though the results have gotten better SKU to SKU.
And you’re quite right they don’t want these chips in the data center and at some point they didn’t really want these cards competing in games with the top end when placed in SLI (when that was a thing) as they removed the connector from the mid range.
Without the extra VRAM, it takes hundreds of times divided by batch size longer due to swapping, or tens of times longer consistently if you run the rest of the model on the CPU.
But if you just want to double the memory without increasing the total memory bandwidth, isn't it a good deal simpler? What's 1 more bit on the address bus for a 256 bit bus?
Instead, it appears entirely possible to double VRAM size (starting from current amounts) while keeping the bus width and bandwidth the same (cf. 4060 Ti 8GB vs. 4060 Ti 16GB). And, since that bandwidth is already much higher than system RAM (e.g. 128-bit GDDR6 at 288 GB/s vs DDR5 at 32-64 GB/s), it seems very useful to do so, though I'd imagine games wouldn't benefit as much as compute would.
You can see this with overclocking VRAM. Greatly benefits mining, slightly or even negatively benefits gaming workloads.
This extends to system RAM too, most applications will see more benefit from better access times rather than higher MT/s.
Works fine with their ollama:rocm docker image on Fedora using podman. No complaints.
Did some gaming, too, just to see how well that works. A few steam games.
Can anyone explain what might be going on here, especially as it relates to power consumption? I thought (bigger die ^ bigger wires -> more current -> higher consumption).
They might be targeting a lower power density per squad mm than compared to amd or nvidia, focusing more on lower power levels.
Instruction set architecture and layout of the chips and PCB also factor into this as well.
I am not a semi expert but bigger die doesn't mean bigger wires if you are referring to cross-section, the wires would be thinner meaning less current. Power is consumed pushing and pulling electrons from the transistor gates which are all of the FET type, field effect transistor. The gate is a capacitor that needs to be charged to open the gate to allow current to flow through the transistor. discharging the gate closes it. That current draw then gets multiplied by a few billion gates so you can see where the load comes from.
All things being equal, a bigger die would result in more power consumption, but the factor you're not considering is the voltage/frequency curve. As you increase the frequency, you also need to up the voltage. However, as you increase voltage, there's diminishing returns to how much you can increase the frequency, so you end up massively increasing power consumption to get minor performance gains.
If Intel is getting similar performance from more transistors that could be caused by extra control logic from a 16-wide core instead of 32.
For tasks that tend to scale well with increased die area, which is often the case for GPUs as they're already focused on massively parallel tasks so laying down more parallel units is a realistic option, running a larger die at lower clocks is often notably more efficient in terms of performance per unit power.
For GPUs generally that's just part of the pricing and cost balance, a larger lower clocked die would be more efficient, but would that really sell for as much as the same die clocked even higher to get peak results?
I should've considered this, I have an RTX A5000. It's a gigantic GA102 die (3090, 3080) that's underclocked to 230W, putting it at roughly 3070 throughput. That's ~15% less performance than a 3090 for a ~35% power reduction. Absolutely nonlinear savings there. Though some of that may have to do with power savings using GDDR6 over GDDR6X.
(I should mention that relative performance estimates are all over the place, by some metrics the A5000 is ~3070, by others it's ~3080.)
Old source, but this says the power cost of increasing the clock frequency is cubic: https://physics.stackexchange.com/questions/34766/how-does-p...
Anyway, expecting good earnings throughout the year as they use Battlemage sales to hide the larger concerns about standing up their foundry (great earnings for the initial 12gb cards, and so on for the inevitable 16/24gb cards).
This strikes me as not a particularly useful metric, or at least one only indirectly related to the stuff that actually matters.
Performance/watt and performance/cost are the only metrics that really matter both to consumer and producer - performance/die size is only used as a metric because die size generally correlates to both of those. But comparing it between different manufacturers and different fabs strikes me as a mistake (although maybe it's just necessary because identifying actual manufacturing costs isn't possible?).
What prevents manufacturers from taking some existing mid/toprange consumer GPU design, and just slapping like 256GB VRAM onto it? (enabling consumers to run big-LLM inference locally).
Would that be useless for some reason? What am I missing?
We will need a new memory design for both GDDR and HBM. And I wont be surprised they are working on it already. But hardware takes time so it will be few more years down the road.
As far as official products, I think the real reason another commentator mentioned is that they don't want to cannibalize their more powerful card sales. I know I'd be interested in a lower powered card with a lot of vram just to get my foot in the door, that is why I bought a RTX 3060 12GB which is unimpressive for gaming but actually had the second most vram available in that generation. Nvidia seem to have noticed this mistake and later released a crappier 8GB version to replace it.
I think if the market reacted to create a product like this to compete with nvidia they'd pretty quickly release something to fit the need, but as it is they don't have too.
Which also answers your question: The manufacturers aren't doing it because they're assholes.
I guess adding memory to some cards is a matter of completely reworking the PCB, not just swapping DRAM chips. From what I can find it has been done, both chip swaps and PCB reworks, it's just not easy to buy.
Software support is of course another consideration.
Once you've filled all the slots your only real option is to do a clamshell setup that will double the VRAM capacity by putting chips on the back of the PCB in the same spot as the ones on the front (for timing reasons the traces all have to be the same length). Clamshell designs then need to figure out how to cool those chips on the back (~1.5-2.5w per module depending on speed and if it's GDDR6/6X/7, meaning you could have up to 40w on the back).
Some basic math puts us at 16 modules for a 512 bit bus (only the 5090, have to go back a decade+ to get the last 512bit bus GPU), 12 with 384bit (4090, 7900xtx), or 8 with 256bit (5080, 4080, 7800xt).
A clamshell 5090 with 2GB modules has a max limit of 64GB, or 96GB with (currently expensive and limited) 3GB modules (you'll be able to buy this at some point as the RTX 6000 Blackwell at stupid prices).
HBM can get you higher amounts, but it's extremely expensive to buy (you're competing against H100s, MI300Xs, etc), supply limited (AI hardware companies are buying all of it and want even more), requires a different memory controller (meaning you'll still have to partially redesign the GPU), and requires expensive packaging to assemble it.
Hardware-wise instead of putting the chips on the PCB surface one would mount an 16-gonal arrangement of perpendicular daughterboards, each containing 2-16 GDDR chips where there would be normally one, with external liquid cooling, power delivery and PCIe control connection.
Then each of the daughterboards would feature a multiplexer with a dual-ported SRAM containing a table where for each memory page it would store the chip number to map it to and it would use it to route requests from the GPU, using the second port to change the mapping from the extra PCIe interface.
API-wise, for each resource you would have N overlays and would have a new operation allowing to switch the resource overlay (which would require a custom driver that properly invalidates caches).
This would depend on the GPU supporting the much higher latency of this setup and providing good enough support for cache flushing and invalidation, as well as deterministic mapping from physical addresses to chip addresses, and the ability to manufacture all this in a reasonably affordable fashion.
GPUs use special DRAM that has much higher bandwidth than the DRAM that's used with CPUs. The main reason they can achieve this higher bandwidth at low cost is that the connection between the GPU and the DRAM chip is point-to-point, very short, and very clean. Today, even clamshell memory configuration is not supported by plugging two memory chips into the same bus, it's supported by having the interface in the GDDR chips internally split into two halves, and each chip can either serve requests using both halves at the same time, or using only one half over twice the time.
You are definitely not passing that link through some kind of daughterboard connector, or a flex cable.
How does "clamshelling" get around the 32-bits per module requirement? Do the two 2GB modules act as one 4GB module when clamshelled?
[1] https://www.reddit.com/r/hardware/comments/182nmmy/special_c...
A non-technical reason is that the market of people wanting to run their personal LLMs at home is very small.
If you harm their profit good luck continuing to have access to GPU chips. It’s a cartel.
As in, a 32bit CPU that runs at 1 giga instruction/second, with a 16 Gbps memory bus, could get up to 0.5 instruction per clock, and that's not very useful. For this reason there can't be an absolute potato with gigantic RAM.
How gigantic is not useful, idk.
Back when they (finally) got into dGPUs seriously, Intel (and everyone else) said it would take many years and the patience to tolerate break-even products and losses while coming up the learning curve. Currently, it looks pretty much impossible to sustain ongoing profitability in low-end GPUs. Given gamer's current performance expectations vs the manufacturing costs to hit those targets, mid-range GPUs ($500-$750) seem like the minimum to have broad enough appeal to be sustainably profitable. Unfortunately, Intel is still probably years away from a competitive mid-range product. Sadly, the market has evolved weirdly such that there's now a minimum performance threshold preventing scaling price/performance down linearly below $300. The problem for Intel is they waited too long to enter the dGPU race, so now this profit gap coincides with no longer having the excess CPU profits to go in the red for years. Instead they squandered billions doing stupid stuff like buying McAfee.
The only really annoying bug I've run into is the one where the system locks up if you go to sleep with more used swap space than free memory, but that one just got fixed.
If you're running Ubuntu, Intel has some exact steps you can follow: https://dgpu-docs.intel.com/driver/client/overview.html
Whoops - included the wrong link! https://www.phoronix.com/review/intel-arc-b580-graphics-linu...
Do you have your config posted somewhere? I'd be interested to compare notes
Intel Arc gpus also support hardware video encoding for the AV1 codec which even the just released Nvidia 50 series still doesn't support.
The newest beta-ish driver is Xe, the main driver is Intel HD, and the old driver is i915.
People complaining experienced the teething issues of early Xe builds.
There's a Mesa DRI driver, called i965 (originally made for Broadwater chipset, thus the 965 numbering), which has since been replaced by either:
- Crocus for anything up to Broadwell (Gen 8)
- Iris for anything from Broadwell and newer
Then there's a Video Acceleration driver, which is (also) called i965. I think this is what you're referring to. There are:
- i965 (aka Intel VAAPI Driver), which supports anything from Westmere (Gen 5) to Coffee Lake (Gen 9.5)
- iHD (aka Intel Media Driver), is a newer one, which supports anything from Broadwell (Gen 8)
- libvpl, an even newer one, which supports anything from Tiger Lake (Gen 12) and up
Battlemage users had to use libvpl until recently because Media Driver 2024Q4 with BMG support was only released 2 weeks ago. Using libvpl with ffmpeg may requires rebuilding ffmpeg, as some distro doesn't have it enabled (due to conflict with legacy Intel Media SDK, so you have to choose).
I have B580 for my Linux machine (6.12), and xe seems pretty stable/performant so far.
iGPU:
| Arch | KMD | DRI (Mesa) | Vulkan (Mesa) | VA |
| -------------------------------------------- | ------- | ----------- | -------------- | ---------- |
| < Broadwater (Gen4) | i915 | i915 | N/A | N/A |
| >= Broadwater (Gen4), < Westmere | i915 | i915 | N/A | i965 |
| >= Westmere (Gen5), < Haswell | i915 | crocus | N/A | i965 |
| >= Haswell (Gen7), < Broadwell | i915 | crocus | hasvk | i965 |
| == Broadwell (Gen8) | i915 | iris/crocus | anv/hasvk | iHD |
| >= Skylake (Gen9), < Tiger Lake | i915 | iris | anv | iHD |
| >= Tiger Lake (Xe/Gen12), < Lunar Lake (Xe2) | i915/xe | iris | anv | iHD/libvpl |
| >= Lunar Lake (Xe2) | xe | iris | anv | iHD/libvpl |
dGPU:
| Arch | KMD | DRI (Mesa) | Vulkan (Mesa) | VA |
| ------------------------------------- | ------- | ----------- | -------------- | ---------- |
| >= DG1 (Xe/Gen12.1), Battlemage (Xe2) | i915/xe | iris | anv | iHD/libvpl |
| >= Battlemage (Xe2) | xe | iris | anv | iHD/libvpl |
Usually, KMD/DRI/Vulkan should work as-is if you use a reasonably recent kernel and mesa, but video acceleration sure is a bit of a mess.Hell, the B580 is CPU bottlenecked on everything that isn't a 7800x3d or 9800x3d which is insane for a low-midrange GPU.
ML stuff can be a pain sometimes because support in pytorch and various other libraries is not as prioritised as CUDA. But I've been able to get llama.cpp working via ollama, which has experimental intel gpu support. Worked fine when I tested it, though I haven't actually used it very much, so don't quote me on it.
For image gen, your best bet is to use sdnext(https://github.com/vladmandic/sdnext), which supports Intel on linux officially, and will automagically install the right pytorch version, and do a bunch of trickery to get libraries that insist on CUDA to work in many of the cases. Though some things are still unsupported due to various libraries still not supporting intel on Linux. Some types of quantization are unavailable for instance. But at least if you have the A770, quantisation for image gen is not as important due to plentyful VRAM, unless you're trying to use the flux models.
My main complaint is that the fan control just doesn't work. They stay at low speed or off no matter how hot the card gets, until it shuts down due to overheating. Apparently there's a firmware update to fix this, but you need Windows to flash it. You can zip-tie a spare fan somewhere pointing at the card...
Secondary complaint is that it's somehow not compatible with Linux's early boot console, so there's no graphical output until the driver is loaded. You'd better have ssh enabled while setting it up.
It's also incompatible with MBR/BIOS boot since it doesn't include an option ROM or whatever is needed to make that work - so I switched to UEFI (which I thought I was already using).
When I ran a shader "competition" some people's code with undefined behaviour ran differently on my GPU than theirs. That's unavoidable regardless of brand and not an Intel thing at all.
On the other hand, I have to keep the driver for the secondary GPU (also intel) blacklisted because last time I tried to use it it was constantly drawing power.
AFAIK it used to be possible to get some SR-IOV working on the previous Alchemists (with some flashing), but Battlemage seems like a proof of Intel abandoning the virtualization/GPU splitting in the consumer space altogether.
If I had more time I'd roll my own site with basic HTML/CSS. It's not even hard, just time consuming.
But yes, platforms usually apply compression in terrible ways, and it's especially noticeable coming from text and straight line stuff like graphs
> According to Intel, the brand is named after the concept of story arcs found in video games. Each generation of Arc is named after character classes sorted by each letter of the Latin alphabet in ascending order. (https://en.wikipedia.org/wiki/Intel_Arc)
(There’s no name decided yet for the fourth in the series.)
> [...] first generation, based on the Xe HPG microarchitecture, codenamed Alchemist (formerly known as DG2). Intel also revealed the code names of future generations under the Arc brand: Battlemage, Celestial and Druid.
https://www.intel.com/content/www/us/en/newsroom/news/introd...
It's fun for the developers and the end-users alike... So no, it's not limited to GPU enthusiasts at all. Everyone likes codenames :)
Except butt-headed astronomers
We were actually told to change our internal names for our servers after someone named an AWS instance “stupid” and I rolled my eyes so hard, one dev ruined the fun for everyone.
So sure, pour one out to whoever's funeral is on the grocery store tabloids that week with your codenames.