Rendering on the Apple M1 Max Chip
blog.yiningkarlli.com
blog.yiningkarlli.com
this is really interesting
With Apple they are just overcharging for solid state drives, which I’m cool with these days (because at least I’ll get them)
Thermodynamics would not take kindly to you having a >99.999...% efficient anything.
Coefficient of performance is a type of efficiency.
> modern heat pumps reach their CoP because they don't actually generate heat, they simply move it around
That's not even true! What a mess of a pedantic correction.
> Thermodynamics would not take kindly to you having a >99.999...% efficient anything
Well the cogen gas powerplants here can produce 50kWh of electricity from burning 100kWh of natural gas. I can use 50kWh of electricity to put 200kWh of heat into my house with a heatpump.
Seems like a good deal to me, and I think carnot would be fine with that.
I know that’s a bummer for people who want to heat their house with their computer, but Thermodynamics is always a bummer, I don’t make the Laws.
I wonder how many CPUs I need to build a cryptocurrency miner water heater... That would be so much better than just wasting energy heating up dumb elements.
Was semi-useful in winter though, all 6 weeks of it...
"If you're plugged in all the time" then try the Mac Pro version when it comes out.
This is about performance in laptop models, where it IS critical.
Except it is still direct electrical heating which is atrociously inefficient.
Electric heating converts practically all energy into heat, making it ~100% efficient. You can make statements about cost-effectiveness compared to burning things, but not all houses can.
CHP configurations are more common in colder climates with district heating, so their "waste" heat during generation often isn't wasted at all.
No, not even close. There are huge losses in electricity production and transmission.
In bigger houses direct electricity just isn't a thing, most have some sort of central heating, and lots have either some combustion or heat pump solution. The latter is gaining.
https://www.anandtech.com/show/17024/apple-m1-max-performanc...
The 5950X is only 55% faster than the M1 Max on multi-threaded integer benchmarks. The M1 Max is even 26% faster than the 5950X on multi-threaded FP benchmarks (maybe it has two AMX units?).
The M1 Max is really in 5900X/5950X territory... in a laptop.
I agree on the pricing (I also have a 5900X and a 5950X machine). But it is fantastic that we can have that kind of performance in a mobile device and still have many hours of battery life.
We should complement both AMD and Apple and be happy that we finally have serious competitors to Intel. AMD has managed to compete with and outrun Intel from the an initial position of an also-ran underdog. I think the M1 line is more impressive technically than Ryzen, given the very low TDPs. But Apple has vastly more budget and much more opportunity for vertical integration.
Both companies have done impressive work the last few years.
For all the dissing of Geekbench, I found it to actually correlate pretty closely with supposedly more elaborate benchmarks like SPEC. (The writing was on the wall for Apple Silicon performance for many years, when the iPhone/iPad Ax CPUs where catching up and surpassing laptop CPUs in Geekbench, but lots of people dismissed it because it was just Geekbench...).
To get some context of the M1 Pro/Max perf, people should take a look at https://browser.geekbench.com/processor-benchmarks, sort by Multi-Core, and slot the M1 Pro/Max in at around 12500.
I would say TSMC's 5nm process is more impressive technically than TSMC's 7nm process, which are used here respectively.
I think Apple will still win on a per Watt basis even when AMD starts using the current-generation process, but the question is: by how much?
Apple bought timed-exclusive access to TSMC 5N and also now 5NP. On the same node the differences would be less.
The comparison should not even be close TBH. We're talking about a laptop chip that can run on battery for an extended period of time that performs as well and sometimes better than the high end consumer AMD desktop chip. Kudos to Apple.
I'm not suprised. Apple has the lead in the fabrication process, 5nm vs 7nm. I eagerly await a true apples to apples comparison when AMD uses the same 5nm process.
The article ends wondering about the impact on PC OEMs. I presume they are extrapolating the performance improvement curves, talking with Microsoft about Windows on ARM and working at HW contingency plans in case they have to leave a sinking x86 ship. I don't think they are resting on promises of big improvements from AMD and Intel.
A 14' Macbook Pro (M1 Max with 32 GPU cores, 32 GB RAM, 512 GB storage) is 3440€.
A PC: 5950X (750€), 32GB RAM (150€), Mainboard (130€), WD SN850 500GB (90€). Now if you build from scratch you need a PSU, CPU cooler and a case (~200€ together).
That's 1320€ without a GPU. Depending on your workload the GPU performance seems to be between a desktop 3060 and 3080. So between 700€ and 1400€.
Tl;dr: A 3440€ Macbook competes with a 2000-2700€ desktop depending on your workload. The desktop has no peripherals and no USB4/TB4.
On the other hand I really want to give the keyboard a shot before I pull the trigger, after 3 years of Thinkpad I think that will be the real pain point...
It's only a little thicker than the macbook pro. It's keyboard doesn't break, and the product line has had a 4k screen since 5 years ago. It's 120Hz refresh rate. It has a very large power brick - 240W, it gets hot and loud with a huge fan exhaust. It's thick metal and about 7-9lb - you can run over it with a car. It only gets 9 hours battery w/ regular usage, and about 3 hours of "fan on time." Keep in mind, with the large and loud fan on, it can stay at 5GHz. This is called a pro laptop - a workstation. When I travel, I bring a 65W PSU, and it runs fine on that, just w/o turbo boost.
No, don't point to the lower "geekbench" score for this laptop - that's not a CPU test. GPU performance is a large part of that test, and they run the test on the default GPU. The M1 only has a single GPU. The Precision's default is the low power integrated graphics, not the discreet GPU. If you have a test where they assign the discreet GPU, please feel free to point it out.
As I've discussed before here, I have a shell script that runs in parallel with a bunch of VMs. My coworkers air (yes, I know it's not the max) runs it in 8-10 hours overnight. I run it over lunch. It loads, does calculations on, and creates graphs from several gig of ascii performance data.
What does compare in performance to the M1 air is my Latitude w/ the I7 in it. The M1 "max" is "max for apple" but competes with mid-tier laptops from everyone else. And that's ignoring the fact that it can't run pretty much any of the useful industry tools or games w/o recompiling x86 to arm on the fly.
Yes, I think my Dell cost the company $7-8k, without support. That's why it's called a pro laptop.
You go on a diatribe about how it doesn’t compete, but your only source is some proprietary workload that you claim is much faster on your laptop.
... but the way you present it undermines your case a bit.
It seems like what you want to say is that if money is no object, weight is no object, heat is no object, battery life is no object, and portability is no object, but the comparison must absolutely be laptop to laptop, then there exists a laptop PC configuration that beats M1. If this is your point, then probably you want to compare an M1 Max to your PC, not your coworker's Macbook Air (which is a fanless laptop...)
I think this is a pretty unusual use case and there aren't too many people who are looking for this exact market segment. I definitely think it's fair to admit that Apple isn't intending to operate in this market segment, for better or worse.
You'd probably also want to drop the part about game support, since anyone who wants to play games can spend 1/4 what your workbench costs and get a shitkicking fast small form factor PC. But also, like, recompilation isn't what you should highlight -- what you should highlight is performance. If the recompilation is fast enough for users not to notice, then it doesn't matter, and if it's not, then the reason why it matters is performance, not recompilation.
Anyway, again, I don't think you're necessarily wrong or whatever here, but you're just presenting your point in a way that I think it's extremely unlikely anyone will care or be convinced.
(Also, "discrete")
The 2TB M1 macbook 16" is $4,300 and the 8TB, max-specced version is $6,100.
I know Apple has a bad rap for high prices, but that machine's prices make Apple prices look bargain basement. You could almost buy TWO M1 macs for the price of one 4TB Dell.
A more practical and better deal would be to get a standard laptop and a beefy workstation/grid you can ssh into.
[Spoiler because I don't think I'll get a response -- there are none. Even when you get into the "luggable" category of workstation that is ostensibly portable but really needs to be plugged in, there is no competition right now. The upcoming Alder Lake should significantly improve Intel's entrant in this category, and hopefully brings some real competition]
But I'm just going to tell it to you now so we don't make the same mistake we have for the past 10 years of computer hardware discussions: specs don't matter. You could tell 90% of the people buying PCs with dGPUs about your 5nm GPU and next-gen power efficiency, but they won't care. They're buying them as gaming devices, general-purpose machines and game development laptops. I'd argue the market for Mac users and PC users has not radically shifted, just the hardware you're using. If we're here to talk smack about hardware superiority, this website would have been insufferable for the past decade, because there was quite literally a complete lack of professional dGPU Macs. Now that the tables have shifted slightly, I don't see why Mac users feel the need to crawl out of the woodwork and declare the game as changed, now that Prometheus gave them the gift of a laptop that doesn't throttle to hell.
I do love this review though:
"It is a nice laptop but extremely noisy. Even when idle the fans are on all the time."
I remember memory bandwidth being the bottleneck when running large(ish) datasets for game worlds. It was so much that we put a lot of pressure on google cloud because we worried they wouldn't be able to compete with bare metal (since it's not usually measured, reported and can be non-guaranteed when you have neighbours).
That doesn't change the fact that the above claim about "laptop Xeon chips" beating the pants off the M1 Max is delusional nonsense.
I have to comment on the RTX 3080 bit: I have used many PC laptops over my career, and currently have a Lenova with a fat, barnburner Nvidia dGPU. The GPU is literally never used, because the moment it engages my battery life falls to cartoonish levels (somewhere in the range of 40 minutes), the laptop becomes a space heater, and the fans turn into jet engines. This is the sort of "spec chasing" that the industry is addicted to, providing absurd, completely unreasonable solutions just so someone can boast. One of the things about Apple, quite contrary to your claim, is that they don't do that. When they provide something, it is meaningfully usable and useful 100% of the time.
Xeon W-11955M (8 cores)
They are what would previously have been branded as "Xeon E3" series chips - on the desktop platform they used to share a socket and be drop-in upgrades with consumer desktop chips, because they're basically the same chips with "enterprise" features like ECC turned on.
An example would be Xeon E3-1285 v3 - which is basically the same thing as an i7 4770.
These products are nowhere, nowhere near the M1 Max. They are consumer laptop chips with ECC and vPro turned on.
I'm also not sure why you're sarcastic about ECC RAM. I have 128GB of RAM in my laptop. If it wasn't ECC, I'd have crashes in my VMs and errors in my calculations. When you go 32GB+ and actually use the RAM, anything that doesn't support ECC cannot be taken seriously for professional use. Like the M1 Max.
So, really no need to test them specifically. Go get an 1185G7 or something and you know what “Mobile Xeon” benches will look like. Anandtech already did those benches.
M1X is using DDR5 which does have on-die ECC.
Apple has never sold a laptop with a Xeon in it.
M1 Max still is the better experience and innovation.
Edit: I know this discussion is partly focused on laptops, but the overall comparison is including server chips so it's clearly not just about laptops. And 60 or 100 watts is child's play in even a tiny desktop.
The 3060 alone is a 60W TGP.
The M1 Max running Cinebench consumes 9W while the 11980HK (previous gen) consumes 45W. Perf/W Is about 3X higher for the M1 Max on CPU alone and I suspect factoring in graphics we’ve got a breakaway here.
Intel will catch up but not today.
I am pretty sure perf/W for the M1 Max chips is excellent, better than the 3060. But saying it performs "on par" with a 3060 isn't really correct.
M1 lovefest won't end just like MacBook lovefest haven't ended all this years when Apple made stuff everyone criticizes now.
Bicycle chains, not so much.
Why?
https://wccftech.com/intel-officially-launches-12th-gen-alde...
As for your power consumption figures, Apple has a supply agreement over 5nm so don't think you can compare architecture apples to apples yet.
The geekbench benchmark is more CPU focused AFAIK, and the article we are discussing is talking about just how much the astonishing M1(.+) memory access bandwidth is improving their rendering time.
The only result I could find with the memory details for the 12900HK didn’t calculate out what the (overclocked) memory bandwidth was: https://gadgettendency.com/intel-core-i9-12900k-lights-up-wi...
I am joking but just a bit. Intel has long history of dirty tricks when it comes to benchmarks - using liquid nitrogen cooling and not mentioning it, manipulating the compiler to generate slower code on competing CPUs, terms of service forbidding publication of benchmarks, fuzzy use (or not mentioning it at all) of TWP/energy usage.
What they say can only be used as an upper bound of how it actually is. The odds are is that it's much worse and with some strings attached as well.
Well be very interesting to see given Alder Lake is dumping AVX-512 which was one of their core performance strengths against AMD. Albeit at staggering cost in thermals and power.
Intel is playing this PR game where they release / announce or leak info, so people might try and wait for what Intel has to offer. It's super obvious, weak, and disingenuous.
Apple doesn't have to score the absolute highest in performance. For better or worse, they need a credible high-end chip that they can marry with their proprietary OS and computer designs to win customers via a compelling package not some benchmark result.
After the M1*, intel is facing an uphill fight to keep other laptop manufacturers on x86. They may well pull together and come out stronger, but the last 10y haven't been stellar and Apple exposed that weakness.
So, will intel win back some of the performance benchmarks? Sure, no doubt. But will Dell win over Apple customers using the latest intel chips? Let's see...
https://www.gettingtechnologyright.com/apples-mac-pro-wheels...
>The second bit of design is in the optional “Afterburner” video editing card. The discussion of it starts at around the 1hr. 26 min. point in the Keynote. It is described as a custom FPGA-based HW accelerator. It is claimed to be capable of processing 6 billion pixels per second. The silicon is not Apple’s but the design is.
Anyway, it’s academical because I don’t see them doing that. Expect maybe to power their own datacenters, but then we’d never hear about it.
But this might a bit in the future as they barely can make enough of them for the MacBooks. I even would assume that the launch of the other computers in the lineup has been put to a later time to be able to keep up at least a bit with demand.
Some are actually hosted in a cloud provider like Google, Microsoft, or AWS.
But they also have a fair amount of their own physical data centers, and I would not be surprised to find that they start filling them out with Apple Silicon hardware.
And they might roll out a lot of services on AWS using Apple Silicon hardware, because AWS has recently made available actual Apple hardware in some of their regions. I think this could potentially be a precursor, but this is just my own personal hunch and I have no information outside my own brain towards this end.
It's not even interesting that the M1 is "slightly" beaten by a Mac Pro's Xeon, because nobody is doing raytracing on CPUs.
A Pascal nvidia GPU will render scenes in Blender faster than a current mid-tier AMD CPU, and it's two generations old.
https://www.phoronix.com/scan.php?page=article&item=blender-...
It takes an RTX 3080 card (not even the fastest consumer Ampere card) 11 seconds to render a scene that the fastest Pascal card (1080) took 94 seconds. That's nearly an order of magnitude faster.
I saw an article comparing the M1 Max to a high end AMD card, but that's also meh. AMDs's GPUs don't perform anywhere near as well as Ampere and they use far more power, so they're an easy target to compare against. AMD fakes their benchmarks by overclocking the cards depending on temperature, and the benchmarks they run to tout performance figures are just shy of where the card starts to throttle back for thermals. An AMD GPU that has been running for an hour will perform substantially worse compared to an Ampere card.
This is backwards -
Nearly all major raytracing workloads today are done primarily on CPUs: vfx, animated film and tv, etc.
There are several reasons for this, including memory limits and the lack of coherence in both path tracing algorithms and common acceleration structures. GPU-based tracing is currently a small subset of rendering tasks (primarily those with real-time requirements).
As GPU memory increases, GPUs will become more attractive in this space - likely in the next few years.
Mental Ray, heavily used in the industry for feature films: GPU accelerated since eleven years ago.
Pixar licensed a bunch of GPU rendering tech from NVIDIA in 2015, so that's six years ago.
Autodesk Arnold
After Effects
Chaos V-Ray
Weta Gazebo and Manuka
...all GPU accelerated for three years or so.
After Effects doesn't really belong on this list?
Gazebo is a preview renderer explicitly designed for GPU use, it's not designed for accurate lighting or material look.
Manuka is not GPU accelerated (it used to be when first written, but it wasn't worth it given complexity of build setup for very little benefit, it's pure CPU now and has been for > 7 years).
Until the GPU renderers can run custom shaders (rather than just the stock shaders the renderers ship with), i.e. OSL supports full GPU support, high end VFX facilities aren't really using GPU renderers that much for final lookdev and lighting.
Smaller facilities which are happy using the stock shaders are though.
You didn't say "feature films", you said "film and TV." You're goalpost-shifting / no-true-scotsman'ing.
> it hasn't been since ~2013 (when Arnold came around).
The Arnold I mentioned that is GPU accelerated? Or a different Arnold? Just checking.
> high end VFX facilities
Again with the goalpost shifting / no-true-scotsman-ing
> aren't really using GPU renderers that much for final lookdev and lighting.
They doing that rendering work on laptops?
What is the article talking about? Server farm hardware?
We're done here.
I haven't moved any goalposts.
Arnold has only been GPU accelerated for a few years - not since 2013. Same with Renderman - XPU only came out this year, and still doesn't support everything on the GPU, including custom shaders (like Arnold GPU doesn't).
Mental Ray really isn't used anywhere. It sort of died like more than a decade ago.
Honestly, I used it back in 2003 in the VFX industry, but not since.
You have to be mistaken here. No one uses Mental Ray that I have heard of in production in the last decade. Probably someone does, but not anyone serious.
Autodesk doesn't even sell licenses of Mental Ray and I believe you will find it hard to buy them from NVIDIA as well: https://www.autodesk.com/products/mental-ray-standalone/over...
Are you thinking iRay? Still made by NVIDIA and uses GPU acceleration.
The default renderer in Maya, 3DS Max is Arnold these days, others use V-Ray, RedShift, Octane, RenderMan, etc....
Surprised no one has mentioned Redshift Renderer yet; GPU-based, supports out of core for both geo and textures, and OSL.
Nevertheless, I agree that "nobody renders on CPU" is just patently false - one need only look at the Corona Renderer numbers to instantly disprove that.
Pixar still primarily renders on the CPU. Only within the last year have they added a hybrid GPU option to the Renderman renderer.
Pixar didn't license any particular special GPU rendering tech from NVIDIA in 2015; they're just using OptiX and CUDA like everyone else. Pixar released initial GPU support in RenderMan this year, but currently XPU only supports a subset of the full RenderMan functionality. Before that, the only GPU ray tracing they were using internally was on an extremely limited basis for early lookdev workflows. Final lookdev and all of lighting continues to be CPU-only. I likely know more about this particular usage than you do because... I worked on the early internal GPU ray tracing for lookdev stuff at Pixar.
Autodesk Arnold has initial GPU support, but it was only added in the past few years. Most studio usage limits Arnold GPU usage to lookdev; most desktop workflows for iterating continue to be CPU because again... not enough memory on GPUs. SPI Arnold has no GPU support at all. Yes, there are two Arnolds out there.
After Effects is not really used for high-end 3D rendering? Why is it on your list? Did you just do a Google search for "GPU renderer" and throw what you found into a list without actually understanding what each thing is?
V-Ray has had extensive GPU support for many years now, and it probably leading the pack among usage in film/vfx/etc houses. However, most film/vfx/etc places use V-Ray CPU for both iterating on the desktop and for final frames, due to GPUs not having enough memory. V-Ray GPU has seen wide adoption in the archviz world though.
Gazebo is an OpenGL viewer, used in DCC viewports. It's not even remotely close to a final frame renderer. Manuka does not have GPU support whatsoever; the Manuka development team literally published an entire ACM TOG paper on this topic which I guess you haven't checked.
I'm not entirely sure how you put your list together; everything on here is either out of date, misunderstood, or just flat out wrong, and all in ways that are immediately apparent to those of us who actually work in this industry.
Nobody is doing realtime raytracing on CPUs. The author is talking about production ray tracing rendering working with scenes that won't fit in most workstation GPUs.
Sadly this is not really interesting, it's disingenuous. He benchmarked against a ten year old (03/06/2012) server CPU. My two year old intel laptop cpu (i7-9750H) also outperforms the xeon's he's comparing against by almost 40%.
The M1 is a great chip, it's sad that this got published with a "server chip" comparison at all. A real server class CPU from the modern era, at a comparable price point, is the AMD EPYC 7443P, which is 400% faster on cpu bench.
Unfortunately I don't exactly have access to piles of fancy server chips sitting around, so I had to make do with just testing against what I have access to, and that's all I had access to. ¯\_(ツ)_/¯
None of them beat an actual, modern, mid-market desktop CPU. And none of those compare to the likes of the threadripper.
The test results are amazing, the exposition here is simply inaccurate. Again, you don't need a "server-class highend workstation" CPU to best the M1 in raw performance. Nothing on the planet beats the m1 in tpd/performance. That's amazing enough on its own.
The remarkable thing about the m1 is the tpd/performance it gets.
Yeah, but for how long? My laptop has a respectable CPU as well but any kind of sustained load would melt it down if it didn't automatically throttle itself back to 2 GHz and below after like 5 seconds.
These M1 chips somehow manage to have great performance while also keeping temperatures and power draw low. I think this will change everything forever.
Bingo
On the other hand I'm sure there's more than a few chipheads out there who are saying "it's about time", there was a longstanding prediction that the arm architecture would overtake x86.
Hat's off to Apple, of course, but Intel were sort of a victim of their own success in that they nearly beat AMD to extinction over the last decade.
I also note that there's no comparison with a properly modern X86 (i.e. an AMD 5000 series). The Apple will still come up on top, most likely, but that Xeon did literally come out in 2012!
Is there really "chipheads" who are predicting ARM ISA to buck this trend and start pulling ahead at equivalent technology nodes? By what mechanism do they believe this will happen, do you know?
I think easier decoders => shorter pipeline, less silicon footprint for decode, possibly more silicon for reordering logic, less dark silicon, etc.
To be fair the chipheads I lived with were gaga about VLIW, so there was certainly a bias there.
If you want to design a DSP or something, where you know it’ll be 100% saturated all the time with an absolutely predictable workload, it’s extremely efficient. Same for AMD Terascale GPUs back in the days before async compute and other GPGPU style tasks took over.
"The theory goes that arm64’s fixed instruction length and relatively simple instructions make implementing extremely wide decoding and execution far more practical for Apple, compared with what Intel and AMD have to do in order to decode x86-64’s variable length, often complex compound instructions."
Not sure it's true, not an expert. But it doesn't sound wrong!
Pre-decode lengths or stop-bits and more recently micro-op caches have been techniques that x86 has used to mitigate this and improve front end widths, for example.
People like Jim Keller (who has actually worked and lead teams implementing these very processors at Apple, Intel, and AMD!) basically say as much (while acknowledging decode is a little harder, in the large scheme of things on modern large cores it's not such a big deal):
https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-...
Andy Glew, one of the architects for Intel's first out of order x86 core (P6) among other things, is another who has said similar.
https://groups.google.com/g/comp.arch/c/elke1FHfYr0/m/SwW9NT...
A consistent 5% win is pretty huge for certain industries.
Are you referring to Andy Glew's thread? He said perhaps 5%, but he also went on to say probably less than 5% for basically the lowest-end out of order processor that was fielded (A-9), not what you would call a high performance core (even then 10 years ago). On today's high performance cores? Not sure, extrapolating naively from what he said would suggest even less. Which is backed up by what Jim Keller says later.
So << 5%, which is significantly less than process node generational increases.
I'm not saying ARM won't leapfrog x86, I'm just asking what the basis is for that belief, and what those who believe it think they know that the likes of Jim Keller does not.
If it's an argument about something other than details of instruction set implementation (e.g., economics or process technology) then that would be something. That is exactly how Intel beat the old minicomputer companies' RISCs despite having "x86 tax", is that they got revenues to fund better process technology and bigger logic design teams. Although that's harder to apply to Apple vs AMD/Intel because x86 PC and server units and revenues are also huge, and TSMC gives at least AMD a common manufacturing base even if Apple is able to pay for some advantage there.
1. The instruction set isn't so much a performance thing as much as a thing that bites you with power usage (you need to have a big fat decoder on all the time in the worst case). The widest X86's can only dispatch 75% of the instruction per cycle but X86 instructions can do more, so you'd have to check a specific benchmark.
2. I want X86 to die more because it's fucking ugly than performance as per se. ARM is not a simple ISA, so although you won't be writing a disassembler in 10 minutes like RISC-V, aarch64 is still much less insane than X86 with all the extensions.
The unwashed masses don't have to care and won't actually use the performance, I'm not that bothered whether they care or not, although it's worth saying that X86 is such a mess that hiding instructions at weird alignments is a valid obfuscation technique - i.e. it's not just aesthetics/performance.
Sadly I have not programmed in assembly since, and I put it down to how ugly the ISA was.
Pretty stupid perhaps, but ¯\_(ツ)_/¯
Eh, it's overhead, but isn't massive overhead. Power usage is much more a function of how the chip was designed.
Does anyone expect x86 to close a factor-of-6 perf/watt difference? (from Anandtech's M1 Max preview) A factor-of-2-to-3 IPC difference? And that's just A15, not against the next-gen A16.
Node makes a big difference, it doesn't close up a factor-of-3 IPC gap in a single node though, that's facially ridiculous. Name a full-node shrink+architectural step that has tripled IPC in the last 10 years. Now name one that has done it while cutting power in 1/6th.
At that point we will see the goalposts shift again and it will be "well, x86 could do it if they wanted but Apple is just more willing to spend more transistors..."
Fact of the matter is the x86 makes it very difficult to spend those transistors efficiently - otherwise it already would have been done. If it was such an obvious gain to just spend those extra transistors, then surely AMD would have done it, if nobody else.
Everyone acknowledges x86 has some problems, but the other thing is that they've already mostly played their hand trying to fix those problems, the known solutions like instruction cache have mostly been exhausted at this point. The idea that AMD and Apple can just triple IPC at a whim but they've chosen not to do so for some reason, is facially ridiculous.
I know what Jim Keller said but the math just doesn't add up on it for me. OK, full node shrink, great, even if that doubles your transistor count at iso-power, or even doubles it at a little less power, that doesn't double your IPC let alone triple it, and it doesn't close a factor-of-6 perf/watt gap.
The one number that surprised me in this review was that the perf/W of Threadripper for the rendering phase: it is very close to the M1 Max. I understand that the numbers are not apples to apples because of the total laptop vs CPU-only comparison, but the power consumption of the Threadripper CPU itself is very high and probably takes the lion's share of the overall power consumption. And that's for a previous generation Threadripper.
it's also a 128-thread processor being put against a 8+2 thread processor, and that's the closest thing to something that will outweigh Apple's IPC advantage here: super wide processor clocked super slow, and unlike the more realistic comparisons (laptop processors, etc) the Epyc has deployed over four times as much silicon just to match the M1.
This is the absolute best-case scenario for x86 - they get six times as much silicon and 16 times as many threads and all they can do is match it.
Do the comparison again against the Mac Pro 40-core chip when it comes out and you'll see A15 pull ahead again.
The 3990x is not designed for energy efficiency, on an older node, and on an older architecture... uses 3kWh vs Apples 2kWh (using a very flawed methodology).
An yet you're claiming Apple has a 3-6x power efficiency advantage.
I'm not sure how that makes any sense.
For a task energy test - which is not the same thing as a perf/watt test - this is completely loading the dice in favor of x86 and M1 still manages to match it. That is an extremely good result for putting a laptop chip against a top of the line HEDT processor in a test that is normally all about race to sleep.
> An yet you're claiming Apple has a 3-6x power efficiency advantage.
I’m not the one claiming anything, if you disagree with Anandtech’s numbers go take it up with them. The numbers don’t change just because you find them uncomfortable.
> In multi-threaded tests, the 11980HK is clearly allowed to go to much higher power levels than the M1 Max, reaching package power levels of 80W, for 105-110W active wall power, significantly more than what the MacBook Pro here is drawing. The performance levels of the M1 Max are significantly higher than the Intel chip here, due to the much better scalability of the cores. The perf/W differences here are 4-6x in favour of the M1 Max, all whilst posting significantly better performance, meaning the perf/W at ISO-perf would be even higher than this.
> In the SPECfp suite, the M1 Max is in its own category of silicon with no comparison in the market. It completely demolishes any laptop contender, showcasing 2.2x performance of the second-best laptop chip. The M1 Max even manages to outperform the 16-core 5950X – a chip whose package power is at 142W, with rest of system even quite above that. It’s an absolutely absurd comparison and a situation we haven’t seen the likes of.
https://www.anandtech.com/show/17024/apple-m1-max-performanc...
And again, that’s with it running at half the clock and half the threads of its laptop peers, so IPC is something like 8x higher in those scenarios.
That is not the kind of gap you close up with a node shrink or tightening pitches a bit. Hackernews experts know best though.
Next year Zen4 and A16 will be on the same node, and then it’ll be another excuse for why x86 is still getting dumpstered. Just keep the goalposts on wheels, you’ll need it.
How do you calculate Perf/watt if not Task/Task Energy?
> As I said, task energy (which is what’s measured in the OP) heavily favors “getting it done quicker”
Which x86 can do.
> It’s basically a “race to sleep”
Oh no, it gets the task done faster!
> I’m not the one claiming anything, if you disagree with Anandtech’s numbers go take it up with them.
I have no issue with Anandtech, since they made no such claim. I read the article, you're badly misquoting it.
> And again, that’s with it running at half the clock and half the threads of its laptop peers, so IPC is something like 8x higher in those scenarios.
None of which is ultimately important.
> That is not the kind of gap you close up with a node shrink or tightening pitches a bit. Hackernews experts know best though.
Look at the results again. For example for SPEC2017 ST the Apple M1 MAX is essentially tied with the 5950x.
Sure the M1 Max might be more energy efficent (though by how much you'd need to measure) - but remember it's a massively larger chip, a whole year newer, and on a newer node.
For MT we see the M1 max bearely beat a 5800X for int, and do significantly better for floating point - which shows different priorities of design.
Again, take a 5800X, shrink it down, add another FPU, and it beats a M1 max hands down.
> Next year Zen4 and A16 will be on the same node, and then it’ll be another excuse for why x86 is still getting dumpstered.
Except x86 isn't getting dumpstered.
You're just cherry picking specific comparisons that make M1 look great, and then ignoring all contrary information.
I mean, I could the the Borderlands 3 1080p benchmark, and say that the M1 got 21.1FPS to the GE76's 100, and therefore x86 is ~8 times faster.
It wouldn't be honest (I'm deliberately chosing the worst M1 chip, and picking a workload that really favors x86) - but I could do it.
But I don't, because it's not honest nor is it helpful.
Anandtech mentions this in their review
https://www.anandtech.com/show/15483/amd-threadripper-3990x-...
Chip Core# TDP 1-Core 1-Core All-Core All-Core Power Freq Power Freq 3990X 64 280 W 10.4 W 4350 3.0 W 3450 3970X 32 280 W 13.0 W 4310 7.0 W 3810 3960X 24 280 W 13.5 W 4400 8.6 W 3950 3950X 16 105 W 18.3 W 4450 7.1 W 3885
As you can see, going from 4.35GHz down to 3.45GHz reduces power consumption by over 3x. Further, per-core power usage of the 3990x is very low overall.
This lower clock and lower per-core performance gives higher overall performance per watt.
I also didn't suggest Apple would never have the best chips ever. Clearly all else being equal if ISA was irrelevant and you had 1 ARM competitor and 1 x86 competitor then sometimes the ARM CPU is going to be the better of the two.
I'm asking is there some continued effect by which people think ARM is going to continue to pull ahead. Is it going to remain < 5%, or is there some turning point where that will start to increase? I'm no expert on this, but there are experts who don't seem to think that there will be such an inflection point.
see my response elsewhere, but these aren't unrelated problems: Apple has higher IPC at a lower power-per-core. You can slide around where on the scale x86 falls - maybe you can match perf/watt but then you're getting wiped by a factor of 3 on performance, and you can match on performance but then you're getting wiped by a factor of 6 on perf-watt. You can't do both at once.
There simply isn't enough transistor gain from a single node shrink there to clear that much of a gap, basically Apple is also seeing much better performance-per-transistor and that's a harder gap to close.
> I'm asking is there some continued effect by which people think ARM is going to continue to pull ahead. Is it going to remain < 5%, or is there some turning point where that will start to increase? I'm no expert on this,
where in the world are you getting that this is <5% gap?
again, 3990WX is an absolute best-case scenario here, that is putting a laptop Apple chip up against a HEDT-class (really, server-class) CPU with 6 times the silicon area and five times the TDP, and all it can do is match it. Mac Pro is the Apple competitor to those chips, and you'll see it slide back into the lead again.
and again, task energy as a measurement favors getting it done faster over pure perf/watt. It's still a 280W TDP / 350W PPT chip against a 60W laptop chip, and it has way more silicon, it's the best case scenario and all they can do is match the M1 in task energy.
That's actually still an extremely good outcome for the M1 and the 40-core Mac Pro is going to slide back over the top again.
I'm not sure how you established that. IPC and picoseconds available per cycle are intrinsically linked. Talking about IPC in isolation is nonsense, particularly when comparing a core that makes less than half the cycle time.
> where in the world are you getting that this is <5% gap?
I just mean the rule of thumb for the "x86 tax", not any specific device. The full thread should have context here, I'm not saying Apple is or is not ahead in a particular instance, I'm asking about more general trends of device performance and ISA.
I’ve already made mine - Anandtech shows a factor of 6 difference in perf/watt between a 11980HK and a M1 Max at peak performance, and this likely translates into a ~factor of four-ish difference in perf/watt and IPC at iso-power. That’s a performance gap that is unlikely to be closed by a node shrink - there is a large architectural gap there. Sure, Apple is probably using tighter pitches as well, but that doesn’t add up to a factor of 4 difference either.
If you have one processor that is doing 4 times the performance at iso clocks, and 2.2x the performance with both processors running at peak clocks, the "megahertz myth" isn't applicable, one of those processors is just faster than the other.
We will see next year, with Zen4 and A16 (apples next core) going head to head on N5P. I strongly doubt Zen4 will even get close.
> I just mean the rule of thumb for the "x86 tax", not any specific device.
Ah, so you are conflating “the amount of transistors spent on x86 decoding” with “the architectural impact that x86 has on performance”.
Unfortunately those are not the same thing. To make the car analogy, how much of your car’s engine bay is spent on aspiration? Probably 5%, maybe 10% right? So obviously aspiration is not important to a car’s performance output at all? And a different method of aspiration would not affect performance at all, a turbocharged car performs almost identically to a naturally aspirated car?
That’s the argument you’re making by focusing on number of transistors spent decoding instead of the impact on the rest of the design. Having a much higher “rate of feed” enables much higher-performance optimizations in the rest of the design - like a much much much deeper reorder buffer.
And just like with cars - that 5% or so of the processor is a key enabling factor that can produce gains of 2x in the rest of the processor, because it’s the only way to keep an engine that is 2x as powerful fed. It doesn’t, itself, produce all that much speedup, but you can’t design bigger engines without clearing that bottleneck. Even if it’s only 5% you can’t do those same designs naturally aspirated.
Similarly, even if the decoder is only 5% of the x86 design, it doesn't mean it's not strangling the ability to scale the rest of the design.
Yes it does, it's refuting your idea that IPC can be considered without looking at cycle time.
And you're trying to make an argument that does not address the question I asked. As I told you in my first reply.
> Ah, so you are conflating “the amount of transistors spent on x86 decoding” with “the architectural impact that x86 has on performance”.
No. Read the conversation from the start, and read the links with comments from people who have actually worked on Apple, AMD, and Intel cores.
Haswell got a very modes (<10%) performance improvement and an equally modest 10% lower power at load (though almost 25% at idle).
Sandy Bridge did much better in performance (up to 40% in some benches) and also did decently well in power consumption despite using the same node.
AMD saw massive increases with Zen on the same 12nm node. The also saw almost 20% IPC increases from Zen 2 to Zen 3 despite both being on the same N7 node
If anything, history shows that node shrinks are always overrated at improving power and performance.
That is irrelevant. What matters is the product of IPC and frequency. x86 parts today are clocked much, much higher than Apple's parts.
IPC and frequency are both means to an end. Compare on performance and efficiency, not implementation details.
The fact of the matter here is that Apple is getting much better IPC at a much better power-per-core, and that is the real architectural gap. You can slide around where on the scale that x86 falls, but there isn't enough gain from a full node shrink to close a factor-of-6 perf/watt gap and a 3x IPC gap.
You could make a very wide CPU indeed if you decided to run it at 100MHz. That would be obviously stupid, though, because it is the product of IPC and frequency that matters.
It certainly looks like Apple has made the better tradeoff. However, you can only tell that from the benchmarks. Either going wider or going faster are valid approaches.
Yes, it’s true that x86 often relaxes pitch, that doesn’t make up for a factor of 6 perf/watt difference like Anandtech measures.
Like I said, I guess we’ll see, Zen4 and A16 will be on the same node next year. By that time the goalpost will move to something else, like this “x86 isn’t designed for power efficiency” defense.
What goalpost have I moved? I have said one (true) thing: that you should judge by performance rather than by implementation details. You are falling victim to the Megahertz Myth, just the other way 'round.
https://images.anandtech.com/graphs/graph17024/117496.png
Why do you think megahertz myth is relevant here? Core for core A15/M1 is plainly faster than any of its peers, ignoring clocks, and it is even farther on top when you do look at clocks (i.e. IPC). It doesn’t matter at all which way you look at it, unless you are putting M1 up against HEDT SKUs like 3990WX - there are a few non-peer scenarios like that it only ties x86 in, like OP looking at task energy (3990WX gets to use 280W TDP/375W PPT and race to sleep) but that’s still an incredibly good outcome considering the loaded test, and Mac Pro with its 32+8 configuration will almost certainly be back on top in the “peer” comparison scenario.
It’s amazing how much breath was wasted on “IPC is what really matters” when Ryzen came out and now it’s “the other side of the megahertz myth” when Apple is on top. Ryzen was never even remotely close to being in the lead on IPC compared to where Apple currently is.
Even at iso-power you are looking at a factor-of-3-to-4 difference in performance - I was being generous with the “only 3x IPC” thing. That is what Anandtech measured in their review. And that still means a gap of 4x perf/watt - which is better than 6x for sure, but it doesn’t mean low-clocked x86 magically beats A15.
IPC is what "really matters" when all the chips you're evaluating are capping out at pretty similar frequencies. When there's a 30% difference in frequency then you need to use instructions per second, and evaluate it with the context of different wattages and different benchmarks.
Finally a sensible comparison! As I said, it is quite impressive.
>it doesn’t mean low-clocked x86 magically beats A15.
Nor did I ever say it would.
What I have said is a simple truth: leading on IPC and trailing on frequency is not obviously an advantage. I don't understand what it is you think I said. Pretty much everything you are writing is a non-sequitur.
In terms of actual raw performance the instruction set is extremely secondary to about 10 other architectural choices.
The reason Apple is getting such performance is because it's caught up with all the cutting edge techniques Intel and AMD use, and a few more, and implemented them inside its core architecture.
These budgets are much tighter on a mobile chip, but are still relevant on desktop design.
So now the wheel turns and someone will have to predict how long it takes before RISC-V overtakes ARM.
There's realistically probably no reason RISC-V cannot work, expect that the number of techniques needed to implement a leading edge high performance evaluator is enormous.
Isn't Mill eternal vaporware?
The ISA has relatively little to do with it. Sure, x86 requires power-hungry decoders, but most of the time you'll be running from the uop cache anyway. Plus you get denser code. Arm and x86 aren't all that different under the hood these days and generally RISC vs CISC is a wash. It took heroic engineering to get x86 that fast, but that work is already done.
Is it? I was under the impression modern CPUs translate x86 instructions do on the fly translation to something else anyway. I’d be very surprised if that hardware translation process was making their chips 4x less power efficient (or whatever the number is).
Also as a consumer, the reasons why x86 stayed around as long as it did sound like a bunch of excuses. If x86 is really as inefficient as you say, then it’s about time the industry started sidelining Intel and moving to a better architecture with or without them. I’m hoping AMD follows suit before long and starts selling competitive arm or risc-v chips.
Windows 10/11 on an M1 Mac via parallels however has been the best Windows ARM device I've ever used and I have no trouble using it daily as my main work laptop.
At 24fps, a 2h movie has 172800 frames. 21,970,310 M1 seconds or 8.5 M1 months. Which is less than I was expecting. The rendering seems to scale linearly per core too.
Presumably bad math or a lot more rendering complexity for the pixar super computer deploys?
But you can denoise a sequence which is similar.
1) that forest isn't near the complexity of most major feature films. It's great, but it's got a ton of instancing, it's not got very complex shaders, it's not got very complex lighting. It's a good benchmark scene, but it's not approaching the level of scenes in most feature films.
2) You're assuming a single render happens per frame. Every scene in a movie is rendered several times. Potentially every single time a new hand touches it, and definitely multiple times as lighters iterate on it. On average I think a single shot must be rendered about 40-50 times or more in a large production. At least 10 times per each major department.
3) Pixar and other studios don't have really extraordinary computers. They're usually run of the mill HP or Dell workstation specced machines with a dual Xeon and tons of RAM.
4) static scenes with diffuse lighting are fast to render. Add in motion and things get tricky. Do you want motion blur in render (as opposed to composited in via motion vectors)? Well now you're sampling multiple frames. What about depth of field? Way more samples to resolve properly. Now add in more lights, and you need more samples per light to resolve a noiseless render.
Actually, a forest with occluding foliage geometry is a quite difficult light transport situation: it's very difficult for next event estimation to work efficiently, so therefore most of the lighting is indirect and therefore is quite noisy, as it's difficult to find the environment / IBL light for the sun.
Portal guiding and NEE guiding can help a bit, but it can still be quite slow.
If only I had enough money to splurge on such curiosity experiments… :-(
For my personal machine, I’d wait for Framework with Ryzen 5000 series…
Good luck with that... it's the one thing that doesn't seem to come up in these comparisons. At least with the Ryzen/Intel CPUs you can easily install Linux - but until this is true of the M1 chips then these benchmarks aren't all that useful.
it would be odd for an article about graphical performance benchmarking to start talking about OS compatibility.
My wife (who while smart doesn’t like tinkering with computers) has been using Ubuntu for two years now and her only complaint is “my hard drive is full and how the heck do I uninstall applications?” Even things like Microsoft Teams and all the virtual doctors meetings apps work for Ubuntu, which surprised me.
I mean Linux distros kinda suck when something goes wrong but I’m honestly impressed how long my wife’s Ubuntu laptop has been working without major issues.
It's a big reason I like Linux.
https://rosenzweig.io/blog/asahi-gpu-part-4.html https://rosenzweig.io/blog/asahi-gpu-part-3.html https://rosenzweig.io/blog/asahi-gpu-part-2.html https://rosenzweig.io/blog/asahi-gpu-part-1.html
https://asahilinux.org/2021/10/progress-report-september-202...
Historically, ray tracing is only used for movies whereas games generally uses rasterization. The former is highly coherent workload, and is great for GPU. The latter is incoherent workload, and generally isn't suitable for GPU.
Games have started using ray tracing, and there have been GPU implementation of ray tracing in the past 5 or so years. What I said above may be stale.
Usually what one does is break the shots down into individual elements rather than render everything as once and then composite it in. Rendering it all in one go can prevent touchups to individual elements. But I am much more VFX than pure animation - I think animation does more one shot one render work.
Funny thing I mostly know what I am talking about.
What studio do you work at?
You are only however talking about your subjective experience and projecting it onto the entirety of the industry.
I've worked in both animated features and VFX at a few of the bigger studios. What sort of studio do you work at?
I'd suggest looking at the landscape of big feature films and seeing what studios are using GPU rendering. It's rare to see any of the big studios using GPU rendering, not just because of existing farm hardware, but because none of the current GPUs and GPU renderers are capable of handling the larger scenes required, as well as losing out on some shader functionality like full OSL support among others...
If you look at the ACM breakdown of production renderers, in their rendering special, there's not a single GPU focused renderer there.
Individual artists may want GPU rendering. However your statement of only old school people using CPU rendering is just ignoring the realities of production.
Wrote and sold software into the VFX industry. Most notable Deadline, a fairly popular render manager. Also wrote/sold another dozen or so tools into the industry.
While CPU renderers are popular among the old crowd, no one wants to use them because they are super slow. The cycle time is brutal and it kills artist productivity. Artists want redshift, octane, GPU cycles, and hybrid Pixar Renderman, basically anything with multi-GPU support. Wherever artists can they want GPU-based renders because they are fast.
Large feature film shops are laggards in adopting new technology and often have a ton of workflows/plugins that can not be easily changed. They have their tired and true pipeline. But up and coming studios tend to be nearly completely GPU-based, because it cuts costs, even if it limits things a bit (but much less than before.)
This is also why you see Unreal Engine getting into architectural rendering -- again because it is a much nicer workflow than waiting around for a few hours for V-Ray to finish up.
Many children TV shows are GPU-rendered now right out of game engines, in part because of the cycle time and the lower visual complexity of the scenes.
So I think we are both right. Large studios are using CPU-based rendering -- you are correct. Artists and more nimble studios are using GPU-based rendering as much as they can because it reduces cycle time while being roughly equivalent costs (well except for the last 2 years because of crypto screwing up GPU prices.)
Remember my original statement you took offence to was:
"3D artists these days want their render boxes to be filled with gpus so they can use cycles or redshift or octane to render fast."
and
"The reason why most artists highly prefer gpus is it gets their cycle time down. Farms often use CPU’s still because of cost and less pressure on a single image render time."
I was speaking about what artists want and I am absolutely correct on that. And what artists want will filter into the rest of the industry -- although a bit slower now because of stupid GPU prices.
You again try and say the big studios are laggard in adopting tech, and have tired pipelines. That's not why they're using CPU rendering. It's because GPU renderers did not scale to meet their needs.
A lot of the big VFX studios are renderer agnostic (e.g ILM), and would have no issue adopting GPU renderers if it met their needs.
You're continuing to project your subjective (and frankly incorrect) opinion that GPU renderers have superseded CPU renderers already onto the industry, without understanding the limitations involved.
You even mention "hybrid Renderman" in which I think you mean XPU, but that's yet another example of a GPU renderer that isn't at parity with the CPU one. The same goes for Arnold GPU etc...
"I was speaking about what artists want" is fine, but you're also dismissing the very real reasons people are still on CPU renderers by saying it's a matter of legacy.
You now say "That is not the part of your statement that I was saying is incorrect."
But your first response was to say that I was "Completely untrue."
Those two statements contradict themselves. That should teach me not to argue on the internet.
Maybe get off the high horse.
As far as I can tell, the biggest problem is simply that GPU raytracing requires a completely different software architecture. Giant boil-the-ocean rewrites that require not only new software but also new hardware are very difficult to justify, especially in a mature industry dominated by (relatively) short-term film production schedules. There are other technical issues too, such as VRAM limits.
Anybody want to take my PhysX card off my hands? :)
Commercials and mid to low budget TV shows are moving to GPU rendering , but it's very much a factor of scene complexity and available time.
Also, these Xeons are really put to shame here.
https://ark.intel.com/content/www/us/en/ark/products/64583/i...
https://ark.intel.com/content/www/us/en/ark/products/193753/...
The MacBook comes with a free display, input devices and a UPS :)
ill just wait for amd 5nm offering that will also be in ram/cpu upgreadeable form factors and Linux compatible out of the box for half the price.
Surely a CPU that is on an older, more tested node (albeit an Intel node, take that as you will) should have amortized the costs of producing silicon products on that node.
The cost of an Apple CPU is cheaper than a worse-performing older CPU. The economics here are surely poor? You can't tell me that Apple of all companies are taking a hit on their new CPUs? I mean, they could be for all I know, but that would be very un-Apple like.
That's a ten year old CPU.
If your old Haswell / Skylake machines still work (and if you don't need Windows 11 which doesn't support them), there's zero technical reason to upgrade to something newer, at least on the desktop.
There's been some progress in terms of power efficiency on laptops, but even there 5+ year old chips are pretty much as good as it gets if you want to stick to Intel. Maybe 10-15% percent slower than the very latest stuff, you won't even notice it. That's why Apple is currently obliterating Intel and AMD - the latter milked their cash cow until it fell over.
And Apple knew that they didn't even need to work particularly hard on core perf - memory bandwidth improvements alone would have delivered the goods. That's another thing people don't get - the CPUs spend quite a bit of their time waiting for bytes to come in or go out. The latencies involved in that are upwards of 200 cycles, and any significant improvement in both the latency and the bandwidth has insane, immediately noticeable benefits for perf, particularly in "modern" languages with poor cache locality. Why Intel / AMD duopoly did so little to address this deficiency is inexplicable, but here we are.
Single core performance on current i9 processors benchmarks at twice the performance of the Haswell chips. Multicore performance is ~4x.
They definitely slept for a while and AMD started taking their lunch money with Ryzen and EPYC but "nothing has really changed since Haswell" is blatantly wrong.
> If your old Haswell / Skylake machines still work (and if you don't need Windows 11 which doesn't support them), there's zero technical reason to upgrade to something newer, at least on the desktop.
Wrong again, mostly because quicksync in that old a processor won't handle current popular streaming video codecs so you end up with non-accelerated video decoding unless you've got a discrete GPU.
> memory bandwidth improvements alone would have delivered the goods.
No advantage from people upgrading from a Haswell's DDR3-1600 (13GB/sec) to Comet Lake's DDR4-2933 (94GB/sec), eh?
Either it's not worth upgrading from Haswell or it is.
For what it's worth, I also threw in comparisons with the 2019 Mac Pro's Xeon W-3245 and the Threadripper 3990X, both of which might be described as "somewhat new".
[link]: https://blog.yiningkarlli.com/content/images/2021/Oct/takua-...
The one you linked looks great to me. Some parts are more convincing than others, but in general I almost feel as if I could touch those trees. Although I admit I don't have a trained eye to see the rendering tricks used.
It doesn't really matter. Nobody's doing rendering on CPUs.
I get that he's sticking to comparing CPU to CPU rendering to compare apples to apples, but that's like saying "My new buggy whip is amazing!" in a news article written in 1950.
Everyone is doing GPU based rendering. Pascal based GPUs are still (I believe) several times faster than even current mid-tier AMD desktop processors. Pascal is two generations behind current GPUs, which are several times faster than it is.
What will actually be relevant and interesting is how the M1X's full capabilities can be leveraged for rendering, and how that will compare against Ampere GPUs.
Everyone in the field is looking at GPU and some significant advances have been made, but it's definitely not true that "nobody" is doing rendering on CPUs.
Is the article about performance in a final render pipeline?
Is this chip sold by apple in a cluster/datacenter-friendly package?
Talk about moving the goalposts.
Does the performance gap close if Intel starts selling similar on-package RAM to consumers?
I suspect yes, and rapidly. They have it, they just apparently don't want to sell it outside of specialized high-margin goods like Xeon Phi.
That assumes on-package RAM is the key to M1 / Apple's SoC performance. And that isn't the case.
https://www.anandtech.com/print/17024/apple-m1-max-performan...
>The M1 max comes with integrated hbm
It is not HBM, just a very wide LPDDR5.
>hence why 3d cache for Zen is a 15% improvement
Increasing Cache size has nothing to do with Bandwidth. It is the latency and cycle count that matters.
>Modern Intel/AMD processors are memory bottlenecked in multicore workloads
Depending on Workload. The whole reason why Apple has put those bandwidth in place was because of GPU, which are bandwidth sensitive. The M1 Pro has the same MT performance as M1 Max despite only have half the memory bandwidth.
You could have much faster memory on an Intel x86 system, and performance wouldn't even make that much different. Your whole CPU uArch needs to be designed take advantage of the additional bandwidth. Longer pipeline with better prefetch.
>and on package memory means the M1 has much higher bandwidth
Again. There are no relationship between on package memory and high memory bandwidth. You could have achieve the same with DDR5 with DIMM slot. At the expense of much higher energy usage.
Stop spreading this myth, please just look at a Vega GPU die shot and you'll know what HBM looks like. You cannot put HBM on a PCB package. HBM is connected with a silicon interposer.
https://images.anandtech.com/graphs/graph15578/115097.png
https://images.anandtech.com/graphs/graph17024/117494.png
You can dig into other benchmarks in the reviews https://www.anandtech.com/print/15578/cloud-clash-amazon-gra...
https://www.anandtech.com/print/17024/apple-m1-max-performan...
They can do whatever they want with 0 compatibility or backward compatibility.
Why is the test not testing against some AMD CPU, CPU that cost $400, 5800x or 5900x for example.
Yes, they can optimize a lot of pipelines but this goes both ways. E.g. they had lightning connector which was good. USB-C beat the pants out of it and now has a much better eco-system. Apple is stuck with their sub-par connector for phones and an unclear strategy (USB-C on macs/iPads but not on phones and low end iPads).
Apple put in quite a bit of effort to make sure x86 binaries still run the ARM Macs. It works for 99.5% of things in my experience. They have even been observed putting in compatibility tweaks for specific applications.
Also fun fact: The x86 compatibility software, Rosetta, isn’t distributed with macOS and has to be downloaded once it’s needed. This is likely Apple hedging their bets against possible future patent claims.
It's difficult to imagine how anyone can think that M1 is "competing" with Intel or AMD because of this. You won't be getting macOS on any new x86 processor, and you won't be getting Windows on an M1.
For the vast majority of people, M1 provides no meaningful decision point. It's not like I'm going to stop using Linux on my servers, pull out of the cloud, and build a closet full of M1 Mac Minis. I could write my day job software from a Mac, but again, I'm not really getting much of a benefit from that because of the rest of the issues with the OS. And gaming is a no-go. So what is the point?
M1 is not a decision point for their consumer base, because most people don't care what CPU architecture is inside their laptop. However, M1 notably improves the decision points that do matter to their customer base: weight, noise, and battery capacity for each customer's workload.
This comes at the cost of macOS and the restrictions inherent in it. If that cost is acceptable, then the benefits are significant. If that cost is unacceptable, then the benefits are irrelevant. As you indicate, that cost isn't acceptable to you, and you do not indicate interest in any of the other benefits of their platform, which specifically exclude Windows, Linux, and PC gaming. This leaves me unable to determine why you're participating in this discussion.
So, then; why does the topic of 'M1 MAX' matter to you at all? I'm happy to continue discussing, but without that information, there's not a lot left to say.
[1] https://en.wikipedia.org/wiki/Porter%27s_five_forces_analysis"But the Linux desktop experience has been ... frustrating to say the least. And I say this as someone who has used linux in the terminal at my job every day for 5+ years now. It's ridiculous that the software experience still cannot match my 9 year old laptop running macOS Sierra on a 2.5GHz Core i5 and 8GB of DDR3 RAM!"
https://wccftech.com/intel-alder-lake-mobility-cpu-benchmark...
So, nothing special about M1, especially if you consider that Adler Lake will use an older manufacturing process (7nm vs 5nm used by Apple).
Not really a victory there. Also, it hasn't even been released yet, so comparing hypothetical unreleased products isn't a good comparison to call released products you can actually buy "nothing special."
Apple never said it compared to a 3080 in Gaming Performance. People made conjectures that it could be roughly in that arena. In Productivity applications (work applications that use a GPU), there it's very close to a 3080.
Also, TFLOPS is a terrible number for performance because it depends on the level of precision and is highly variable on workload. Good luck loading, say, a 48GB 3D Model on a RTX 3090 - but a M1 Max would handle it fine. Similarly, the 10.8 TLOPs PlayStation 5 had better performance than the 12 TFLOPs Xbox Series X for months after launch.
Finally, as for many of the games that were compared - notice how many of them use MoltenVK, which is a Vulkan -> Metal translator that impacts performance. Then look at how many were x86, needing Rosetta translation on top of that. These weren't exactly the most native of game ports.
“Nothing special” compared to an unreleased Intel chip with higher power usage.
“Overstated by fake numbers” by choosing 1 metric and ignoring 99 others it’s better at.
You’re also a 2 hour account.
Is M1 vs Intel personal to you?
How many pounds of battery would an Alder Lake laptop required to match the uptime of the M1 MAX laptop performing those rendering tasks?
My mental napkin math suggests extremely unfavorable answers for an Alder Lake laptop, to the degree that the laptop's viability in the market would be severely compromised by sustaining the power and cooling required by Alder Lake for long-duration workloads such as rendering.