Raytracing on Intel's Arc B580
chipsandcheese.com
chipsandcheese.com
> Intel uses a software-managed scoreboard to handle dependencies for long latency instructions.
Interesting! I've seen this in compute accelerators before, but both AMD and Nvidia manage their long-latency dependency tracking in hardware so it's interesting to see a major GPU vendor taking this approach. Looking more into it, it looks like the interface their `send`/`sendc` instruction exposes is basically the same interface that the PE would use to talk to the NOC: rather than having some high-level e.g. load instruction that hardware then translates to "send a read-request to the dcache, and when it comes back increment this scoreboard slot", the ISA lets/makes the compiler state that all directly. Good for fine control of the hardware, bad if the compiler isn't able to make inferences that the hardware would (e.g. based on runtime data), but then good again if you really want to minimize area and so wouldn't have that fancy logic in the pipeline anyways.
This is incorrect for AMD, which has "s_waitcnt" instructions in its ISA, which is publicly documented. I believe it is also incorrect for Nvidia, but don't have the receipts to prove it.
I'm also hoping that Intel puts out an Arc A770 class upgrade in their B-series line-up.
My workstation and my kids' playroom gaming computer both have A770's, and they've been really amazing for the price I paid, $269 and $190. My triple screen racing sim has an RX 7900 GRE ($499), and of the three the GRE has surprisingly been the least consistently stable (e.g. driver timeouts, crashes).
Granted, I came into the new Intel GPU game after they'd gone through 2 solid years of driver quality hell, but I've been really pleased with Intel's uncharacteristic focus and pace of improvement in both the hardware and especially the software. I really hope they keep it up.
You get me equivalent of a $500 Nvidia card for around $300 or less. And it makes sense because Intel knows if they can get a foothold in this market they're that much more valuable to shareholders.
Great for gaming, no real downsides imo.
And then just watch heads explode.
I wonder if they could mix clamshell mode and quadrank to connect 64 memory chips to a GPU. If they connected 128GB of VRAM to a GPU, I would expect then to sell it for $2000, not $600.
GPU should be about $200 at TSMC (400-450mm2).
+ about $150 for the pcb, cooler and other stuff, I didn't consider
Times a 1.6 to 1.75 factor if they like actually being profitable (operations, rnd, sales, marketing, ...).
So about $1.5k, I guess.
Multiply that with a .33 "screw the competition" factor and my initial guess is almost spot on.
.
Real problem:
The largest GDDR7 package money can buy right now is 3GB. That's a 1376bit bus right there. GL fitting that to a sub 500mm2 die.
In the future you could put that amount of vram on a 512bit bus, tho.
Also normal DDR is getting really fast atm. 8 channel can already challenge most vram configurations. Maybe it's time soon to switch back to swappable memory.
Assuming I had access to gerbers I could order replica of 5090 PCB for $65, including shipping. Intel PCB is half that. Again this is for a dude off the street buying 1-5 copies, not a bulk order.
What they should do instead is make the cards thinner and more efficient so that you can easily put two of them in a case.
Even selling at cost is a subsidy.
I'm proud to support them. Intel is also selling their lunar lake chips fairly cheaply too. Let's all hope they make it through this rough patch. I can't imagine a world where we only have one x86 manufacturer.
Does it even matter? Some people won’t notice even if there are zero x86 manufacturers.
In fact I would say lots of people have not bought x86 CPU in while, between Mac, RPi and risc-v boards…
Competition is always good
A number of businesses have switched to using arm EC2 servers from x86 EC2 servers for lower costs and things work fine on them.
x86 isn't going anywhere for a long time.
Imagine if AMD never existed. Everything about owning a computer would be worse. Likewise now we need Intel ( or someone else, perhaps a Chinese OEM) to make competition.
> Acer Arc A770 16gb GPU w/Free Game & Shipping $229.99 $229.99 $399.99 at Newegg
Sure that margin is holding when they had to mark the first generation down to get them off the shelves. It would truly surprise me if they've made a significant profit off these cards.
Intel are "fools" for not adding a feature that maybe a few thousand people care about?
For reference the B580 die is nearly the size of the 4070 but sells for a third the price.
Doesn't this suggest the B580 has worse yields? Die surface area isn't directly proportional to selling price.
What is says is that given the die area, Intel fails to capitalize on their chip relative to a 4070.
On the die size argument, which I see being echoed a lot online:
Why would a customer care or factor that into their purchasing decisions? Saying that these dGPUs with large dies are what is going to put Intel out of business is ludicrous and the Xe cores are shared amongst many of Intel's most lucrative products.
You can afford larger dies on N4 compared to when the 40-series were launched. It's no longer the leading edge node and yields have likely improved.
dGPUs have pretty expensive GDDR modules, I do not have data on the exact proportion but I would bet that the memory modules is the more important line item.
BoM matters less on lower volume (compared to mobile SoCs) dGPU units. Masks masks, R&D and validation are big fixed up-front costs.
Recurring software support is also independent of how many units get sold. Xe cores are shared by many Intel products (client & server CPUs, datacenter GPUs and gaming GPUs).
B580 is widely popular for gamers, Intel cannot keep up the demand at the moment. I doubt they need to unlock SRIOV on the gaming segment dGPUs to get rid of stock, as you seem to suggest. Their datacenter GPUs [1] offer support for SRIOV, as you probably already know, so I assume you are bemoaning market segmentation.
[1] https://www.youtube.com/watch?v=tLK_i-TQ3kQ -- Wendell's video on Flex 170 GPU from Intel - Subscription Free GPU Accelerated VDI on Proxmox 8.1
Of course consumers don't care about die size and cost, only the value of the end product. The problem is intel and their negative margins. Maybe people are too blitz scale brained but running a negative margin hardware business is a really bad idea. Typically they target 50% margins.
Now if intel was healthy maybe they could survive a few negative margin products by subsidizing from their profitable SKUs. However intel is not healthy and reported losses every quarter last year. It's a sinking ship and they need a come to Jesus monent before they go bankrupt.
I'm arguing that their product segmentation strategy is stuck in the past and a significant reason why they're unprofitable. On desktop CPUs they lock ECC even though the Xeons are cheaper because they remove the E cores so AVX512 will work. On dGPUs they lock out SRIOV even though the desktop iGPUs support it. On the server they created accelerators which they then tried to charge for which means no one bothered to write integrations to make it work with their software. It's a broken culture which is so up it's own ass that management only cares about intel and is completely disconnected from the customers. Intel needs product features and value differentiators that aren't just negative margin products, changing product segmentation is low hanging fruit.
Being intransigent and the same as the competition but slightly worse had led them to unprofitability and soon, bankrupty. Interest rates are going up and companies carrying debt are going to be in a rough spot.
They need to stop artificially crippling their products to segment them. Simplify naming and have mid, great and greatest. The lowend shit products should be dumped. This means SRV-IO and ECC available everywhere. All of their GPUs should shift right by 8GB or more.
Streamlining their SKUs would streamline their wafer processing in the fabs.
They do have too many SKUs though.
At this point they should just produce Xeons and support at most 2 sockets. I don't care if they fuse off features with cryptographic keys. Ship one piece of silicon to everyone, streamline the shit out of that.
ECC everywhere, no questions asked.
I haven't checked in a while but for client there's basically two die masks. In 13th and 14th gen half the SKUs were 12th gen masks.
Mobile is either the same as client or a smaller mobile efficiency die, but only one mobile mask.
For server there's just E cores and P core tiles. The limitation is packaging not fab.
The SKU is just fuses and written at validation after binning.
ECC everywhere would be a far better excuse for enterprises to upgrade their old PCs that 'AI PCs'
AMD Ryzen CPUs have ECC enabled but not officially supported. Intel still locks away the feature.
What I have assumed given then trend, but could be completely wrong about, is that the raytracing version of the world might be easier on the software & game dev side to get great visual results without the overhead of meticulous engineering, use, and composition of different lighting systems, shader effects, etc.
It's ironic that you harp about "hacks" that are used in rasterization, when raytracing is so computationally intensive that you need layers upon layers of performance hacks to get decent performance. The raytraced results needs to be denoised because not enough rays are used. The output of that needs to be supersampled (because you need to render at low resolution to get acceptable performance), and then on top of all of that you need to hallu^W extrapolate frames to hit high frame rates.
You can see the purely ray traced part in this image from the post: https://substack-post-media.s3.amazonaws.com/public/images/8...
This combination of techniques is actually pretty smart: Combine the powers of the rasterization and ray tracing algorithms to achieve the best quality/speed combination.
The rendering implementation in software like Blender can afford to be primitive in comparison: It's not for real-time animation, so they don't make use of rasterization at all and do not even use denoising. That's why rendering a simple scene takes seconds in Blender to converge but only milliseconds in modern games.
For primary visibility, you don't need more than 1 sample. All it is is a simple "send ray from camera, stop on first hit, done". No monte carlo needed, no noise.
On recent hardware, for some scenes, I've heard of primary visibility being faster to raytrace than rasterize.
The main reasons why games are currently using raster for primary visibility:
1. They already have a raster pipeline in their engine, have special geometry paths that only work in raster (e.g. Nanite), or want to support GPUs without any raytracing capability and need to ship a raster pipeline anyways, and so might as well just use raster for primary visibility. 2. Acceleration structure building and memory usage is a big, unsolved problem at the moment. Unlike with raster, there aren't existing solutions like LODs, streaming, compression, frustum/occlusion culling, etc to keep memory and computation costs down. Not to mention that updating acceleration structures every time something moves or deforms is a really big cost. So games are using low-resolution "proxy" meshes for raytracing lighting, and using their existing high-resolution meshes for rasterization of primary visibility. You can then apply your low(relative) quality lighting to your high quality visibility and get a good overall image.
Nvidia's recent extensions and blackwell hardware are changing the calculus though. Their partitioned TLAS extension lowers the acceleration structure build cost when moving objects around, their BLAS extension allows for LOD/streaming solutions to keep memory usage down as well as cheaper deformation for things like skinned meshes since you don't have to rebuild the entire BLAS, and blackwell has special compression for BLAS clusters to further reduce memory usage. I expect more games in the ~near future (remember games take 4+ years of development, and they have to account for people on low-end and older hardware) to move to raytracing primary visibility, and ditching raster entirely.
I will admit that I was a bit sly in that I omitted the word "realtime" from the path tracing part of my claim on purpose. The amount of denoising that is currently required doesn't excite me either, from a theoretical purity standpoint. My sincere hope is that there is still a feasible path to a much higher ray count (maybe ~100x) and much less denoising.
But that is really the allure of path tracing: a basic implementation is at the same time much simpler and more principled than any rasterization based approximation of global illumination can ever be.
However, that's not what's driving raytracing.
The vast majority of game development is "content pipeline" - i.e. churning out lots of stuff - and engine and graphics tech is built around removing roadblocks to that content pipeline, rather than presenting the graphics card with an efficient set of draw commands. e.g. LoDs demand artists spend extra time building the same model multiple times; precomputed lighting demands the level designer wait longer between iterations. That goes against the content pipeline.
Raytracing is Nvidia promising game and engine developers that they can just forget about lighting and delegate that entirely to the GPU at run time, at the cost of running like garbage on anything that isn't Nvidia. It's entirely impractical[1] to fully raytrace a game at runtime, but that doesn't matter if people are paying $$$ for roided out space heater graphics cards just for slightly nicer lighting.
[0] That one scene in The Stanley Parable notwithstanding
[1] Unless you happen to have a game that takes place entirely in a hall of mirrors
For the artists, being able to wiggle lights around all over in real time was an immeasurable productivity boost over even just 10s of seconds between baked lighting iterations. They had a selection of options at their fingertips and used dynamic lighting almost all the time.
But, that came with a lot of restrictions and limitations that make the game look dated by today’s standards.
(I mean, for games that are mostly static. I can definitely see why some games might want to be raytraced because they want some dynamic stuff, but that isn’t every game).
One of the effects I really like is bounce lighting. Especially with proper color. If I point my flashlight at a red wall, it should bathe the room in red light. Can be especially used for great effect in horror games.
I was playing Tokyo Xtreme Racer with ray tracing, and the car's headlights are light sources too (especially when you flash a rival to start a race). My red car will also bounce lighting on the walls in tunnels to make things red.
It doesn't even have to be super dynamic either, I can't even think of a game that has opening a door to the outside sun to change the lighting in a room with indirect lighting (without ray tracing it). Something I do every day in real life. It would be possible to bake that too, assuming your door only has 2 positions.
Research and progress is necessary, Ray tracing is a clear advancement.
AMD could just easily skip it if they want to reduce costs, we could just not by the gpus. Non of it is happening.
It does look better and it would be a lot easier if we would only do ray tracing
Secondly, nvidia are a company that want to sell stuff for a high asking price, and once a certain tech gets good enough that becomes more difficult. If the 20 series was just a incremental improvement from the 10, and so on then I expect sales would have plateaued especially if game requirements don't move much.
Screenspace Ambient Occlusion? Marching rays (tracing) against the depth buffer to calculate a terrible but decent looking approximation of light occlusion. Some of the modern SSAO implementations like GTAO need to be denoised by TAA.
Screenspace Reflections? Marching rays against the depth buffer and taking samples from the screen to generate light samples. Often needs denoising too.
Light shafts? Marching rays through the shadow map and approximating back scattering from whether the shadowed light is occluded or not.
That voxel cone tracing thing UE4 never really ended up shipping? Tracing's in the name, you're just tracing cones instead of rays through a super reduced quality version of the scene.
Material and light behavior is not the problem. Those are constantly being researched too, but the changes are more subtle. The big problem is light transport. Rasterization can't solve that, it's fundamentally the wrong tool for the job. Rasterization is just a cheap approximation for shooting primary rays out of the camera into a scene. You can't bounce light with rasterization.
For rasterization to be useful it must approximately do the same thing that light does in the real world. Therefore rasterization that wants to get closer and closer to the real world will have to emulate more and more of the real world.
It will have to cast exactly the rays that rasterization is hoping to avoid.
Well, I am a gamedev, and currently lead of a rendering team. The answer is very simple - because ray tracing can produce much better outcomes than rasterization with lower load on the teams that produce content. There's not much else to it, no grand conspiracy - if the hardware was fast enough 20 years ago to do this everyone would be doing it this way already because it just gives you better outcomes. No nvidiabux necessary.
> There's not much else to it, no grand conspiracy
True, in that raytracing is the future. Though I don't think it's a conspiracy rather than just the truth that "RTX" as a product was Nvidia creating a 'new thing' to push AMD out of. Moat building, plain and simple. Nvidia's cards were better at it unsurprisingly, much like mesh shaders they basically wrote the API standard to match their hardware.
And just to make sure Nvidia doesn't get more credit than it deserves, the debut RTX cards (RTX 20 series) were a complete joke. A terrible product generation offering no performance gains over the 10 series at the same price with none of the cards really being fast enough to actually do RT very well. They were still better at RT than AMD though so mission accomplished I guess.
Another example is when was the last time you've seen a game with a mirror that wasn't broken?
Working mirrors are limited to less complex scenes in GTA. Hitman too I believe.
See:
>It's something that's trotted out as a nice gimmick, 99% of the time it's not there, and you don't really notice that it's missing.
Yeah, it's a nice detail for the 1% of time that you're in a bathroom or whatever, but it's not like the immersion takes a hit when it's missing. Moreover because the game is third person, you can't even accurately judge whether you'll be spotted through a mirror or not.
simplified tldr: with raytracing you build the environment, designate which parts (like sun, lamps) emit light and you are done. With regular an artist has to spend hours to days adding many fake lightsources to get same result.
What does RTX do, what does it replace, and what does it enable, for whom? Repeat for Physx, etc. Give yourself a bonus point if you've ever heard of Nvidia Omniverse before right now.
It required discarding a lot of "tricks" that had been learnt with rasterisation to speed things up over the years, and made things slower in some cases, but meant everything could use raytracing to compute visibility / occlusion, rather than having shadow maps, irradiance caches, pointcloud SSS caches, which simplified workflows greatly and allowed high-fidelity light transport simulations of things like volume scattering in difficult mediums like water/glass and hair (i.e. TRRT lobes), where rasterisation is very difficult to get the medium transitions and LT correct.
Not sure how much RDNA 4 and on will improve it.
Path tracing is a specific technique where you ray trace multiple bounces to compute lighting.
In recent games, "ray tracing" often means just using ray tracing for direct light shadows instead of shadow maps, raytraced ambient occlusion instead of screenspace AO, or raytraced 1-bounce of specular indirect lighting instead of screenspace reflections. "Path traced" often means raytraced direct lighting + 1-bounce of indirect lighting + a radiance cache to approximate multiple bounces. No game does _actual_ path tracing because it's prohibitively expensive.
The other significant part is that path tracing is independent of the number of light sources, which isn't the case for some of the classical ray traced effects you mention ("direct shadows" vs path traced direct lighting).
That's at least what I understand of the matter.
I didn't manage to get it for MSRP (because living in Europe does tend to increase the price quite a bit, a regular RTX 3060 is over 300 EUR here), but I have to say that it's a pretty nice card, when most others seem quite overpriced or outside of my budget.
When paired with an 5800X the performance is good, the XeSS upscaling looks prettier than FSR and pretty close to DLSS, the framegen also seems to have higher quality than FSR (but more latency, from what I've seen), the hardware AV1 encoder is lovely and the other QSV ones are great, though I do wish that I could get a case big enough and a new PSU to have both A580 and B580 in the same computer and use the B580 for games and A580 for the other stuff (not quite sure how well that combination would work, if at all).
Either way, I'm happy that I got the card, especially with a decent CPU (even the A series with my previous Ryzen 5 4500 was an absolute mess, no software showed the CPU being maxed out but it very much was a bottleneck) and do kind of hope that I'll get the likes of performance that you get in War Thunder, or even GTA V Enhanced Edition for the years to come (yes, the raytracing works there as well) or even more recent games like Kingdom Come: Deliverance 2.
If the upscaling/framegen support was even better in most game engines and games, then it could be stretched further or at least used as a band aid for the likes of Delta Force or Forever Winter - games that come out with pretty bad optimization and are taxing on the hardware, with no good way to turn subjectively unnecessary effects or graphical features off, despite the underlying engines themselves being able to scale way down.
At the end of the day, even if Intel Arc won't displace any of the big players in the market, it should improve the market competitiveness which is good for the consumer.
If they started shipping GPUs with more RAM, I think they'd be in a strong position. The traditional disruption is to eat the low-end and move up.
Silly as it may sound, but a Battlemage where one can just plug in DIMMs, with some high total limit for RAM, would be the ultimate for developers who just want to test / debug LLMs locally.
Reminds me of this old satire video: https://www.youtube.com/watch?v=s13iFPSyKdQ
Making some kind of super niche card for a small fraction of developers who might want the option to upgrade RAM, just doesn’t make financial sense. It’ll end up being more expensive for everyone compared to just buying the card with the amount of RAM you need soldered in.
I mean, check the Framework Desktop. If it doesn’t even make sense for Framework, who can justify a bit of extra cost for modularity, it doesn’t make sense for any company.
Until you do the math.
128GB of the slowest possible DDR4 in 8 DIMMs is still faster than 16GB of the fastest DDR5 in one DIMM. 12.8GB/s*8 = 102.4GB/s > 74GB/s
It requires many more traces, but you still get more throughput overall with more, cheaper memory further away.
The packaging costs on the GPU do go up for the extra traces.
More critically, no one would design it like that. There's a memory hierarchy. You would likely be a small amount of high-speed soldered-on RAM next to the GPU, and more slightly further out.
If Intel could give away a Battlemage to every PyTorch developer, every game developers, and every deep learning developer, even tossing in a check for free cash, in return for it becoming their primary video card, the ROI would be astronomical.
If Intel gave _me_ a Battlemage, and I actually used it (I wouldn't; I have an NVidia), the expected ROI would likely be >$1M. For myself alone.
The key gap between AMD's $170B market cap and NVidia's $2.92T market cap is software and ecosystem.
Making a card specifically for developers makes a ton of sense, since what those developers develop will then have many, many orders of magnitude more __users__.