This is relevant. This means that the performance increase vs Intel is using the extremely throttled 1.2 GHz i7 as the baseline.
This is relevant. This means that the performance increase vs Intel is using the extremely throttled 1.2 GHz i7 as the baseline.
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
Quote:
"In the overall SPEC2006 chart, the A14 is performing absolutely fantastic, taking the lead in absolute performance only falling short of AMD’s recent Ryzen 5000 series.
The fact that Apple is able to achieve this in a total device power consumption of 5W including the SoC, DRAM, and regulators, versus +21W (1185G7) and 49W (5950X) package power figures, without DRAM or regulation, is absolutely mind-blowing."
"The craziest thing is how close [to 1/56th of a Xeon 8176] they get at 3W. That's the part that's staggering. What could Apple do with, say, laptop thermals? Desktop thermals? The mind creeps closer to boggling." -- me, 5 October 2018
Guess we're finding out.
That said, Apple won't make their chips available to the greater market. They might use variations of it in their own data centers, maybe.
I wonder if this could potentially be a good time to bring back Xserve...
1) Is Apple interested (and capable) making macOS into a realistic choice for servers again?
2) If not, is Apple comfortable with selling a box that most people will put Linux on? Would they put resources into a Bootcamp-esque driver package?
"See, you don't need native Linux on the machine, just use our hypervisor & Docker/VMs"
Experience counts, but Apple is pretty old-hat at this by now as well.
I think it's a good CPU, but I don't think it'll be a great CPU. Judging by how heavily they leaned on the efficiency, I am pretty sure that it will be just good enough to not be noticeable to most Mac users.
Can't wait to see actual benchmarks, these are interesting times.
The important part there is "in their class".
I'm sure the Apple silicon will impress in the future, but there's a reason they've only switched their lowest power laptops to M1 at launch.
The higher end laptops are still being sold with Intel CPUs.
They said the Macbook Air (specifically) is "3x faster than the best selling PC laptop in its class" and that its "faster than 98% of PC laptops sold in the last year".
There was no "in its class" designation on the 98% figure. If they're taken at their word, its among every PC laptop sold in the past year, period.
Frankly, given what we saw today, and the leaked A14x benchmarks a few days ago (which may be this M1 chip or a different, lower power chip for the upcoming iPad Pro, either way); there is almost no chance that the 16" MBPs still being sold with Intel chips will be able to match the 13". They probably could have released a 16" model today with the M1 and it would still be an upgrade. But, they're probably holding back and waiting for a better graphics solution in an upcoming M1x-like chip.
The most powerful Macbook Pro, with a Core i9-9980HK, posts a Geekbench 5 of 1096/6869. The A12z in the iPad Pro posts a 1120/4648. This is a relatively fair comparison because both of these chips were released in ~2018-2019; Apple was winning in single-core performance at least a year ago, at a lower TDP, with no fan.
The A14, released this year, posts a 1584/4181 @ 6 watts. This is, frankly, incomprehensible. The most powerful single core mark ever submitted to Geekbench is the brand-spanking-new Ryzen 9 5950X; a 1627/15427 @ 105 watts & $800. Apple is close to matching the best AMD has, on a DESKTOP, at 5% the power draw, and with passive cooling.
We need to wait for M1 benchmarks, but this is an architecture that the PC market needs to be scared of. There is no evidence that they aren't capable of scaling multicore performance when provided a higher power envelope, especially given how freakin low the A14's TDP is already. What of the power envelope of the 16" Macbook Pro? If they can put 8 cores in the MBA, will they do 16 in the MBP16? God forbid, 24? Zen 3 is the only other architecture that approaches A14 performance, and it demands 10x the power to do it.
For example, I can say with 100% confidence that M1 has nowhere near 32MB of onboard cache. Once it starts hitting cache limits, it's performance will drop off a cliff as fast cores that can't be fed are just slow cores. It's also worth noting that around 30% of the AMD power budget is just Infinity Fabric (hypertransport 4.0). When things get wide and you have to manage that much wider complexity, the resulting control circuitry has a big effect on power consumption too.
All that said, I do wonder how much of a part the ISA plays here.
Another important consideration is the on-SOC DRAM. This is really incomparable to anything else on the market, x86 or ARM, so its hard to say how this will impact performance, but it may help alleviate the need for a larger cache.
I think its pretty clear that Apple has something special here when we're quibbling about the cache and power draw per core differences of a 10 watt chip versus a 100 watt one; its missing the bigger picture that Apple did this at 10 watts. They're so far beyond their own class, and the next two above it, that we're frantically trying to explain it as anything except alien technology by drawing comparisons to chips which require power supplies the size of sixteen iPhones. Even if they were just short of mobile i9 performance (they're not), this would still be a massive feat of engineering worthy of an upgrade.
The bigger issue here is bandwidth. AMD hasn't increased their APU graphics much because the slow DDR4 128-bit bus isn't sufficient (let alone when the CPU and GPU are both needing to use that bandwidth).
I also didn't mention PCIe lanes. They are notoriously power hungry and that higher TDP chip not only has way more, but also has PCIe 4 lanes which have twice the bandwidth and a big increase in power consumption (why they stuck with PCIe 3 on mobile).
It's also notable that even equal cache sizes are not created equal. Lowering the latency requires more sophisticated designs which also use more power.
So you expect the M1 MBP will outpace the more expensive Intel 13 MBP they are selling for more money!!
How would that not destroy sales of their “higher end” intel MBP 13?
If so do you think it will be better across most workload types?
Some people especially developers may be skeptical of leaving x86 at this stage. I think the smart ones would just delay a laptop purchase until ARM is proven with docker and other developer workflows.
Another consideration - companies buying Apple machines will likely stay on Intel for a longer time, as supporting both Intel and ARM from an enterprise IT perspective just sounds like a PITA.
I run Linux on all my machines, and even running many (5-10) containers, 16GB was plenty. I now understand a bit better.
I would not buy a new machine today for work with less than 32 GB.
There are likely many people who are not ready to switch yet either.
It has been a while since the ppc/x86 transition, but I want to say it was a similar situation then
This release is entirely within Apples control, why would they risk damaging their brand releasing a chip with lower performance than the current Intel chips they are shipping. They would only do this at a time when they would completely dominate the competition.
But, just looking at A14 performance and extrapolating its big/little 2/4 cores to M1's 4/4; In the shortest tldr possible; Yes.
M1 should have stronger single-core CPU performance than any Mac Apple currently sells, including the Mac Pro. I think Apple's statement that they've produced the "world's fastest CPU core" is overall a valid statement to make, just from the info we independent third-parties have, but only because AMD Zen 3 is so new. Essentially no third parties have Zen 3, Apple probably doesn't for comparison, but just going on the information we know about Zen 3 and M1, its very likely that Zen 3 will trade blows in single core perf with the Firestorm cores in A14/M1. Likely very workload dependent, and it'll be difficult to say who is faster; they're both real marvels of technology.
Multicore is harder to make any definitive conclusions about.
The real issue in comparison before we get M1 samples is that its a big/little 4/4. If we agree that Firestorm is god-powerful, then can say pretty accurately say that its faster than any other four-core CPU (there are no four-core Zen 3 CPUs yet). There's other tertiary factors of course, but I think its safe enough; so that covers the Intel MBP13. Apple has never had an issue cannibalizing their own sales, so I don't think they really care if Intel MBP13 sales drop.
But, the Intel MBP16 runs 6 & 8 core processors, and trying to theorycraft what performance the Icestorm cores in M1 will contribute gets difficult. My gut says that M1 w/ active cooling will outperform the six core i7 in every way, but will trade blows with the eight core i9. A major part of this is that the MBP16 still runs on 9th gen Intel chips. Another part is that cooling the i7/i9 has always been problematic, and those things hit a thermal limit under sustained load (then again, maybe the M1 will as well even with the fan, we'll see).
But, also to be clear: Apple is not putting the M1 in the MBP16. Most likely, they'll be revving it similar to how they do A14/A14x; think M1/M1x. This will probably come with more cores and a more powerful GPU, not to mention more memory, so I think the M1 and i9 comparisons, while interesting, are purely academic. They've got the thermal envelope to put more Firestorm cores inside this hypothetical M1x, and in that scenario, Intel has nothing that compares.
Anandtech's spec2006 benchmarks of the A14 [0] suggest the little cores are 1/3 of the performance of the big ones on integer, and 1/4 on floating point. (It was closer to 1/4 and 1/5 for the A13.) If that trend holds for the M1's cores, then that might help your estimates.
[0] https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
The meme of saying that Geekbench is not a useful metric across cores, or that it is not representative of real-world usage, and therefore cannot be used needs to die. It’s not perfect, it can never be perfect. But it’s not like it will be randomly off by a factor of two. I’ve been running extremely compute-bound workloads on both Intel and Apple’s chips for quite a while and these chips are 100% what Geekbench says about them. Yes, they are just that good.
In this case, it's high speed differential signaling, and that's going to have a /lot/ of active power. There's a lot of C*dv/dt going on there!
To me it seems pretty unlikely to be that important because if you can have 16gb of memory in the chip, how hard can it be to increase the caches a fraction of that?
Next, SRAM doesn't scale like normal transistors. TSMC N7 cells are 0.027 nanometers while MY cells are 0.021 (1.35x). meanwhile, normal transistors got a 1.85x shrink.
I-cache is also different across architectures. x86 uses 15-20% less instruction memory for the same program (on average). This means for the same size cache that x86 can store more code and have a higher hit rate.
The next issue is latency. Larger cache sizes mean larger latencies. AMD and Intel have both used 64kb L1 and then move back to 32kb because of latencies. The fact that x86 chips get such good performance with a fraction of the L1 cache points more to since kind of inefficiency in Apples design. I'd guess AMD/Intel have much better prefetcher designs.
No, when your chip has many billions of transistors that’s not a big number. For 1 mb that’s about 0.2%, a tiny number, also when multiplied with 1.85.
Next the argument is that x64 chips are better because they have less cache while before the Apple chips couldn’t compete because Intel had more. That doesn’t make sense. And how are you drawing conclusions on the design and performance of a chip that’s not even on the market yet anyway?
Maybe I'm misunderstanding, but the 1.85x number does not apply to SRAM.
I've said for a long time that the x86 ISA has a bigger impact on chip design and performance (esp per watt) than Intel or AMD would like to admit. You'll not find an over-the-top fan here.
My point is that x86 can do more with less cache than aarch64. If you're interested, RISC-V with compact instructions enabled (100% of production implementations to my knowledge) is around 15% more dense than x86 and around 30-35% more dense than aarch64.
This cache usage matters because of all the downsides of needing larger cache and because cache density is scaling at a fraction of normal transistors.
Anandtech puts A14 between Intel and AMD for int performance and less than both in float performance. The fact that Intel and AMD fare so well while Apple has over 6x the cache means they're doing something very fancy and efficient to make up the difference (though I'd still hypothesize that if Apple did similar optimizations, it would still wind up being more efficient due to using a better ISA).
I’ll just wait for the independent test results.
If 32mb would be too hard they could have easily went for 1mb.
But they didn’t and that’s a pretty good indicator it doesn’t make a lot of difference.
>but this is an architecture that the PC market needs to be scared of.
This just means you know nothing about processor performance or benchmarks. If it was that easy to increase performance by a factor of 3x why hasn't AMD done so? Why did they only manage a 20% increase in IPC instead of your predicted 200%?
[0] https://images.anandtech.com/doci/14892/a12-fvcurve_575px.pn...
The closest it could get I think would be running a variant of Unix optimized for the Ryzen.
There's no way to say what "faster" means or what 98% they used. Is it faster than 98/100 models? or faster than 98% of the 261M laptops sold in 2019?
"faster than 98% of laptops sold last year" is a nice quotable soundbyte that will be spread without any of those details.
I'm actually a big fan of Apple's products. I'm sure all the numbers, R&D, and benchmarking will prove the M1 to be impressive.
Keep in mind that the $3k MBP is probably part of the 2% in the above quote. The large majority of laptops sold are ~$1k, and not the super high end machines.
If you buy a Ferrari, they won't market saying its latest car is better than 98% of cars sold last year.
it is not a good metric to compare performance of budget laptops where sales is going to be higher to premium laptop.
Plus, given use of Rosetta 2 they probably need 2x or more improvement over existing models to be viable for existing x86 software. Interesting to speculate what the M chip in the 16" will look like - convert efficency cores to performance cores?
The M1 supposedly has a 10w TDP (at least in the MBA; it may be speced higher in the MBP13). If that's the case, there's a ton of power envelope headroom to scale to more cores, given the i9 9980HK in the current MBP16 is speced at 45 watts.
I'm very scared of this architecture once it gets up to Mac Pro levels of power envelope. If it doesn't scale, then it doesn't scale, but assuming it does this is so far beyond Xeon/Zen 3 performance it'd be unfair to even compare them.
This is the effect of focusing first on efficiency, not raw power. Intel and AMD never did this; its why they lost horribly in mobile. Their bread and butter is desktops and servers, where it doesn't matter. But, long term, it does; higher efficiency means you can pack more transistors into the same die without melting them. And its far easier to scale a 10 watt chip up to use 50 watts than it is to do the opposite.
My only worry about the systems with more cores (Mac Pro etc) are about the economics for Apple of making these chips in such small volume.
PS Interesting that from Anandtech the M1 has a smaller die area than the i5/i7 in the Intel Airs so plenty of room for more cores!
If you want a more efficient processor you can just reduce the frequency. You can't do that in the other direction. If your processor wasn't designed for 4Ghz+ then you can't clock it that high, so the real challenge is making the highest clocked CPU. AMD and Intel care a lot about efficiency and they use efficiency improvements to increase clock speeds and add more cores just like everyone else. What you are talking about is like semiconductor 101. It's so obvious nobody has to talk about it. If you think this is a competitive edge then you should read up more about this industry.
>Their bread and butter is desktops and servers, where it doesn't matter.
Efficiency matters a lot in the server and desktop market. Higher efficiency means more cores and a higher frequency.
>But, long term, it does; higher efficiency means you can pack more transistors into the same die without melting them.
No shit? Semiconductor 101??
>And its far easier to scale a 10 watt chip up to use 50 watts than it is to do the opposite.
You mean... like those 64 core EPYC server processors that AMD has been producing for years...? Aren't you lacking a little in imagination?
Smacking that with single cores on a 5nm process is probably pretty easy.
I'm really impressed by what they've done in this part of the market - which would have been impossible with Intel CPUs.
AKA, I suspect a big part of those perf uplifts evaporate when compared with the laptops that are appearing in stores now.
And the ipad benchmarks also remain a wait and see, because the perf uplifts I've seen either are either safari based, or synthetic benchmarks like geekbench. Both of which seem to heavily favor apple/arm cores when compared with more standard benchmarks. Apples perf claims have been dubious for years, particularly back when the mac was ppc. They would optimize or cherry pick some special case which allowed them to make claims like "8x" faster than the fastest PC, which never reflected overall machine perf.
I want to see how fast it does with the gcc parts of spec..
I saw some graphs that looked unrealistically smooth with two unlabeled axis
And I got so excited about it...
When the name i7 is mentioned people are more likely to think it's one of the high end desktop processors rather than the mobile processors.
But then they throw in code compilation... that's where they stop being honest.
Not Intel's fastest chip but, if these benchmarks are correct, that's a big speed bump and the higher end Pro's are to come. How fast does it have to be in this thermal envelope to get excited about?
"Apple claims the M1 to be the fastest CPU in the world. Given our data on the A14, beating all of Intel’s designs, and just falling short of AMD’s newest Zen3 chips – a higher clocked Firestorm above 3GHz, the 50% larger L2 cache, and an unleashed TDP, we can certainly believe Apple and the M1 to be able to achieve that claim."
Not sure what the custom instructions you're referring to are.
M1
> Apple M1 chip
> 8-core CPU with 4 performance cores and 4 efficiency cores
> 8-core GPU
> 16-core Neural Engine
Intel
> 1.7GHz quad-core Intel Core i7, Turbo Boost up to 4.5GHz, with 128MB of eDRAM
https://www.apple.com/macbook-pro-13/specs/
https://www.apple.com/shop/product/G0W42LL/A/refurbished-133...
Apple's claim:
> With an 8‑core CPU and 8‑core GPU, M1 on MacBook Pro delivers up to 2.8x faster CPU performance¹ and up to 5x faster graphics² than the previous generation.
https://www.apple.com/shop/buy-mac/macbook-pro/13-inch-space...
Fine print:
> Testing conducted by Apple in October 2020 using preproduction 13‑inch MacBook Pro systems with Apple M1 chip, as well as production 1.7GHz quad‑core Intel Core i7‑based 13‑inch MacBook Pro systems, all configured with 16GB RAM and 2TB SSD. Open source project built with prerelease Xcode 12.2 with Apple Clang 12.0.0, Ninja 1.10.0.git, and CMake 3.16.5. Performance tests are conducted using specific computer systems and reflect the approximate performance of MacBook Pro.
> Testing conducted by Apple in October 2020 using preproduction 13‑inch MacBook Pro systems with Apple M1 chip, as well as production 1.7GHz quad‑core Intel Core i7‑based 13‑inch MacBook Pro systems with Intel Iris Plus Graphics 645, all configured with 16GB RAM and 2TB SSD. Tested with prerelease Final Cut Pro 10.5 using a 10‑second project with Apple ProRes 422 video at 3840x2160 resolution and 30 frames per second. Performance tests are conducted using specific computer systems and reflect the approximate performance of MacBook Pro.
A more fair comparison would be Tiger Lake (20% IPC improvement) on Intel's terrible 10nm process. The most fair comparison would be zen 3 on 7nm, but even that is still a whole node behind.
I'm not sure exactly what your point is. The fact it is IT IS on a better node. That's probably a big part of the performance picture. But, as a user I don't care if the improvement in performance is architectural or node. I just care that it's better than the competition.
So comparing a
4-Core, 5nm, 16 MB cache, 10 watts TDP vs
4-Core, 14nm, 8 MB cache, 15 watts TDP.
https://ark.intel.com/content/www/us/en/ark/products/192996/...
Apple claims the M1 to be the fastest CPU in the world. Given our data on the A14, beating all of Intel’s designs, and just falling short of AMD’s newest Zen3 chips – a higher clocked Firestorm above 3GHz, the 50% larger L2 cache, and an unleashed TDP, we can certainly believe Apple and the M1 to be able to achieve that claim.
Thats per core performance.
The "Neural Engine" is a special compute unit specifically for machine learning. The iPhone Camera app uses it for semantic segmentation, for example.
It's all the wires and buffers and pipeline registers and arbiters and decoders and muxes that connect the components together.
It's dominated by wires, typically, which is probably how it came to be known as a fabric. (Wild speculation on my part.) I've been hearing that term for a long time. Maybe 20 years?