As was already noted in other comments, the M2 Ultra bandwidth is not that special against high-end GPUs (all the recent ones have over 700GB/s generally) and this bandwidth has to be shared with CPU. So technically if you keep doing work on the CPU there is less available bandwidth at any given time. A 4090 + 13900K has almost 1100GB/s combined ; not that it matters for most use case.
For regular CPU tasks the added bandwidth doesn't seem to make a difference as far as I can tell, at least Apple Silicon isn't winning in any scenario where they don't have a specialized block on the chip for the task. So what's the point ? (beside overpaying for memory)
And to "win" this the total VRAM available at once is what make a difference, not really the bandwidth ; this is just because the task has been parallelized as much as possible. Even then, it required to optimize for the arch with parallelization and it is absolutely not cost competitive with the PC used as a reference. If you really need to maximize GPU VRAM in a single workstation (without going to a servers/cloud solutions) you could build a machine with multiple RTX A4000 SFF (1 slot, 20GB). It would get more expensive than the maxed out M2 Ultra but at this point the M2 Ultra lose so hard in FLOPS power that you really need to specifically look for situations where you would want more VRAM (up to 144GB available for the M2 Ultra GPUs vs 80GB for 4x1 slot card) but wouldn't want to run the model faster/longer in a dedicated server rack that could potentially have even more available VRAM (and be available/shared with other peoples).
Realistically NVidia knows how to put more RAM in their GPUs because it doesn't make sense to scale VRAM faster than computer power for most workload, you need to have a balance that make sense.
As an analogy it is like coming up with a truck than can carry 150T at once but can only do so at 1/3rd the speed of regular trucks. In most case you actually gonna want to run 3 regular trucks even thought is going to be less efficient (it still gonna cost less and be faster overall) unless you really don't have a choice ; at this point you are in "special convoy" territory (like for wind turbine blades) and it's gonna cause lots of headaches on top of being slow and expensive.
Apple market their stuff as an incredible innovation when in fact not only it is irrelevant for most workload that are usually thrown at workstations (mobile or not) but I would argue that running the workloads where it would actually make a difference is a bad idea on a single user workstation. For most things that actually matters in a single user workstation/prosumer/enthusiast system, Apple Silicon lose quite hard especially when it comes to GPU performance : viewport performance, close to real-time 3D rendering (before sending to render farm for final detailed render), games, etc...
And this is the Ultra version of the chip, that is out of reach for most people (it makes look at the 4090 as not that overpriced, which is quite funny). If you go down to the M2 Max version, suddenly the bandwidth is 400GB/s and not only it is not impressive at all, it is even worse than an Intel A770M laptop GPU (512GB/s) while still having less raw power and costing way more. The more you go down in the Apple Silicon roster, the worse it gets. AS is not competitive at the high-end workstation level but it is absurdly overpriced at almost every level.
The reason they have this architecture (that isn't very good for most traditional computer application) isn't because they went out of their way to engineer something great. Nope. It is because they basically scaled up a mobile architecture that was like this from the get go (power and space constraint, plus no need to have that much RAM nor have it upgradeable). And this is only because Apple is currently run by a Scrooge who figured he could get even more money out their silicon division if they solds SKUs with binned parts and controlled the RAM supply/price.
If Apple had actually done useful engineering they would have figured out a way to scale the GPU/VRAM combo independently and a way to package/sell it efficiently. It makes no sense to scale VRAM past a certain point : why would you want to load a 3D model/view/whatever if you cannot compute it fast enough. As for the CPUs existing memory interfaces where fast enough for most things and the "benefit" is inexistent in most case.
They went about it in the worse way possible with cost reduction above all approach while jacking up the price up to 11. This is the most lazy approach they could take and they even dumped all the unnecessary cost directly onto the consumer (low yield for big area chips and soldered RAM close to the chip from a lack a dedicated GPU SKUs). Even if the consumer want to absorb the cost he still get bad scaling and uncompetitive performance...
I just don't get how Apple get away with it and there are people like you falling for their marketing bullshit that is just a spin on what are actually weaknesses...