And yeah, as others point out, this is Apples to oranges. x86 desktops are great at some things, M2 Ultras are great at others, and the overlap that really matters is pretty small... Like, you have to be crazy to buy an M SoC for gaming, or buy a Nvidia GPU for workloads that won't fit in VRAM.
I imagine that in the future, something like Steam will wrap this functionality to provide the ability to run the whole library under the toolkit. And individually-published games will do the same so they install and run with a more consumer-friendly experience.
I have seen the demos, but I am skeptical of the actual practicality or value proposition until a 3rd party publishes some frametime benchmarks, and games out in the wild get battle tested.
(There is the Neural Engine which supports lower precision, but it limited in various ways.)
Regardless, the strides that Apple has been making are impressive.
Although quantisation has now lowered the required memory, I wouldn’t be surprised if it comes in handy again in the near future.
But today that’s quite a niche use case.
Then there's the power consumption difference to consider. This seems like one of those cases where benchmarks reveal only a fraction of the larger picture.
Apple have an opportunity, if they 2-4x the memory on the entry level devices (not beyond the realms of possibility), to make local inference a thing available to all.
A lot of work is going on in 8-bit inference and even 4 bit inference. So, models that need 64 GB in FP32 can do with 16GB VRAM in FP8 or INT8, which is well within the realm of consumer NVIDIA cards. And the latest NVIDIA tensor cores will absolutely destroy Apple Silicon GPUs or the Neural Engine in 8 bit.
So, I don’t think it’s really a strong argument. And as someone who is a Mac user and a ML practitioner, I’d be very happy if they started supporting eGPUs again.
Apple Silicon has many strengths and the GPU core are fine for many ends, from games to graphics apps.
But let’s not pretend that Apple is beating NVIDIA at their own game (yet). That day might come, but currently it only leads to disappointed users in ML forums who were hyped into thinking that their vanilla M2 MacBook Airs can almost compete with a 4090 in training a deep transformer model. (Yes, that happens.)
https://wccftech.com/m2-ultra-only-10-percent-slower-than-rt...
Computation per kWh (or rate of computation per kW) is the right efficiency metric, not TDP (thermal design power).
But it won't happen before all M2/3 Macs will be obsolete. Having goals like that is great but until we reach that point efficiency is still important.
What would be more interesting is to see how Nvidia's laptop cards fare here though - they're constrained to much lower wattage (80-120w) and would make for a much fairer fight against the ~200w M2 Ultra.
Doesn't look like Apple offers Ultra in a laptop - just the Basic, Pro, and Max.
That being said, it's pretty obvious that Apple's mobile-style solution isn't really working out on the desktop side of things. The new iMac feels starkly pedestrian compared to the old ones, and the Mac Mini/Studio are both neat but not unprecedented. The M2 Ultra represents a lot of engineering effort going into flipping that status quo, but its still slipping behind by a considerable margin. Don't forget that a second "Ultra" style SOC with 4x M1 Maxes was supposedly cancelled for drawing too much power and being too hot. It's just not effective or efficient to force that much silicon that close together.
Then you’d be looking for a 24 core CPU, 64GB RAM, 1TB PCIe 5 SSD, mainboard with 6x thunderbolt ports, a silent cooling setup, high quality case that is both small and all the gear while running cool. and if you’re stuck with MS Windows - an Operating System.
It doesn’t matter if the cooling solution gets the chops’s surface temperature lower if it’s heating the room twice as fast.
Chip surface temperature is not a useful metric for this purpose.
If Apple clocked their GPU to match the performance of a comparable dedicated chip, it would be just as inefficient, noisy and hot. Except they can not do even do that. They turned a limitation of the design into a supposed feature.
Source?
https://twitter.com/0xDEADBEEFCAFE/status/166747612998729728...
Regardless, wccftech is far from reliable. IIRC, /r/amd blocks links to the site.
Those are nvidia's best consumer GPUs. I think the cheese grater falls into the pro segment. In that segment nvidia has the A6000s with 48GB VRAM and 91 SP TFlops compared to the 4090's 24GB and 73 SP TFlops. But that costs as much as the Mac Pro alone. And even bigger options (segmented for server/datacenter use) are available.
power consumption doesn't scale linearly with performances
the absolute best of class NVIDIA discrete GPU offering could possibly outperform the Apple GPU at the same power level
Or, to put it in another way, to recover that remaining 50% of performances (2x) the increase in power consumption would be exponential (a lot more than 2x, like 10x)
As far as I understand it—and this is just from watching Apple's presentations on the architecture—the lack of a discrete GPU is a big part of how the Apple Silicon machines achieve good performance per watt.
Instead of having discrete RAM or a discrete GPU with its own VRAM, all of the RAM is accessible to the CPU and and the GPU in a unified memory architecture. On the M2 Ultra, this allows for 800 GB/s of memory bandwidth, and also eliminates a lot of the need to copy data from RAM to VRAM, as both the GPU and the CPU can access the same memory. In return, this allows the GPU to match the performance of discrete GPUs that have a lot more cores.
Of course, the big downside is that you can't expand the RAM or install a beefier GPU. It's all baked in to the logic board.
Plenty of PC hardware reviewers have done sensitivity analysis experiments to see how discrete GPU performance is affected by running with a slower or narrower PCIe link. The consensus is usually that GPUs connected by PCIe have more than sufficient bandwidth, and cutting it in half only affects gaming framerates by a few percent. Tighter coupling between CPU and GPU can plausibly have a bigger impact for some GPU compute workloads, but for traditional 3D graphics it doesn't help performance much.
If once a minute you miss 4 frames in a row that’s noticeable even if everything else is rock solid. The thing is people adjust their resolution/settings to reach acceptable FPS, thus it’s fairly GPU independent. What matters is rendering volatility as assets are loaded etc.
All adaptive refresh rates do is avoiding unwanted delays after the frame finishes.
Adaptive refresh rate means that only frame rate matters. You can't miss a frame, you can only have a frame take longer, and that's measured perfectly well by 1% or 0.1% low frame rates.
Also I don't see why an M series CPU is going to be any better at feeding the (asynchrous, multiple frames in flight) GPU rendering pipeline than a modern x86 CPU, when both are just as fast. macOS is also pretty bad at realtime scheduling for heavy workloads compared to modern Linux and Windows.
A 10 minute test at 120 FPS is 10 * 60 * 120 ~= 72,000 frames. Your 1% lows is an average of ~720 frames but so what if of 300 them are at 60 FPS you’re going to have trouble noticing.
While if you’re almost rock solid at 120FPS a dozen 70ms stalls it’s really obvious that something is wrong but the calculated 1% low’s may actually look better. This is especially true as those stalls generally correlate with interning things happening.
The M series CPU comparison people are talking about isn’t average FPS or even 1% lows where dedicated graphics cars have an advantage, the comparison is what happens when things go very wrong. And it’s in that very specific case when they may have an advantage.
I think this is an instance of the coordinated omission problem: you cannot simply measure the latency of the frames that were completed, but instead have to consider the frames that should have been delivered during a stall. Looking at frame time percentiles means a stall only penalizes your metric with one bad frame, when it should be penalized for missing several frames and showing the user an increasingly stale image for several average-frametimes.
Think of it like Little's law in queing theory , it's better to simply just look at different percentiles, like 99 or 99.9th. The average framerate correctly indicates the average staleness.
It's not clear to me: are you saying those two scenarios should be quantified as equally bad? Because it seems pretty obvious to me that the 5x outlier is qualitatively much worse.
Benchmarking sites use 1% or 0.1% because they better highlight the commodity PC hardware rather than the games bugs and architecture.
These benchmarks are controlled for game bugs and architecture by simply averaging over ~40-50 games, on repeatable loops.
The only advantage is that the transfer from GPU to CPU is faster in terms of latency. This doesn't cause pipeline stalls, because as I've explained above, frames are kept in flight, so latency is not critical. At the same time, modern CPU-GPU interconnects have similar bandwidth to RAM.
Additionally, you're not the first person to have thought about frame pacing. Dozens of reviewers have full frame pacing graphs with the frame times of every frame, as well as 0.1% lows, and we simply don't see what you're describing. 100ms+ frames are not a problem on modern games and modern hardware, and when they are, it's because of some blocking read to storage or to some scheduling issue, not because of CPU/GPU speed.
The impact of link speed on 0.1% low framerates has already been investigated, and it's minimal, and there isn't even an advantage here.
There’s a multitude of such bugs which ship with modern titles. Benchmarking sites don’t use the release day build of 2077 for very good reasons.
It enables software to be designed differently ie by not having to ever copy things to VRAM.
While 192GB of RAM is more than I would need, for people looking to use 1.5TB of RAM and a pair of NVIDIA GPU they had from previous model, they’d have to go elsewhere.
Which leaves me wondering; how much engineering at Apple are happening on Mac?
The best case this article can make is that if you need to play the latest game or do intense ML stuff you probably want NVIDIA, but that's the same as it ever was.
I don't think it was ever meant to be the most informative article: it seems written to serve the contrarians because that's profitable from a readership perspective for a publication like this.
A 4080 is best of class?
And second of all, for the price it should be compared to the 4090, which absolutely demolishes it.
Second, people don't buy Macs only for performance. They also buy Macs for macOS, for integration between devices, for a system that is cool and quiet, for hardware acceleration of ProRes, for on-device privacy-preserving machine learning. Being a bit slower than competing AMD and Intel systems is acceptable, because you get so many other desirable properties in return.
I'd definitely consider a Mac Studio with an M2 Max or M2 Ultra, if I didn't want something portable. I would never buy a machine with a competing machine with an Intel or AMD machine, because I don't want to deal with Windows or desktop Linux.
Other people have another set of priorities and that is fine.
While I’m still surprised they didn’t put a second Ultra in the Mac Pro, I’m betting there’s a wider delta than people imagine between the two form factors.