M1 Icestorm cores can still perform well
eclecticlight.co
eclecticlight.co
let vC = [Float](repeating: 1, count: 4)
func unidiomatic() -> Float {
var vA = [Float](repeating: 0, count: 4)
let vB = [1, 2, 3, 4] as [Float]
var tempA: Float = 0.0
for _ in 0..<1000 {
for i in 0...3 {
tempA += vA[i] * vB[i]
}
for i in 0...3 {
vA[i] = vA[i] + vC[i]
}
}
return tempA
}
func idiomatic() -> Float {
var vA = [Float](repeating: 0, count: 4)
let vB = [1, 2, 3, 4] as [Float]
var tempA: Float = 0.0
for _ in 0..<1000 {
tempA = zip(vA, vB).map(*).reduce(0, +)
for (index, value) in vA.enumerated() {
vA[index] = value + vC[index]
}
}
return tempA
}
When compiled with optimizations (it's unclear if the author did this?), the "idiomatic" code here is indeed much worse. It's easy to fix the worst of that by working on a lazy sequence, i.e. tempA = zip(vA, vB).lazy.map(*).reduce(0, +)
which will save on an extra array allocation and instead run the sum "in place" (in this case, vB is recognized as a constant and in both cases and distributed accordingly). Unfortunately there is still a small amount of overhead because zip doesn't know that the arrays are of the same length, so the compiler inserts some extra bounds checks; also, the second loop gets transformed into what is essentially temp = __swift_intrisic_move(vA)
for (index, value) in vA.enumerated() {
temp[index] = value + vC[index]
}
vA = __swift_intrisic_move(temp)
but this should ideally bring the performance of the two much closer.Also, just to be entirely clear, an example like this (once I “fixed” it) is probably the best Swift can reasonably expect to get, since it hits all the hardcoded semantic hints that Array has been annotated with. It actually does get quite close to the unidiomatic version; save for some extra bounds checks from the zip and a dubious double move, it’s basically done as well as it could with constant propagation and inlining. More complex examples are unlikely to fare as well.
(This can't happen automatically because the default semantics generally involve a panic or exception to be triggered when the check fails; there's no way to optimize this further.)
Transforming from reduce to map-reduce is usually very trivial, but has to be kept in mind
```
swiftc -emit-sil ...
```
Then if that doesn't workout, then try to dump LLVM IR:
```
swiftc -emit-ir ...
```
Both compiler flags should be available in out-of-box Swift compiler.
https://www.cpubenchmark.net/compare/AMD-Ryzen-7-PRO-5850U-v...
but is 19% slower in single core performance. However, if you consider that AMD uses 7nm and Apple 5nm technology to build their processors, AMD is a lot better.
I'd love to see actual measurements for both chips.
[1] https://www.notebookcheck.net/AMD-Ryzen-7-PRO-5850U-Processo...
Ryzen : https://www.cpubenchmark.net/cpu.php?cpu=AMD+Ryzen+7+PRO+585...
M1 : https://www.cpubenchmark.net/cpu.php?cpu=Apple+M1+8+Core+320...
TDPs don't mean anything even within a vendor lineup, and doing cross comparison is futile (without properly measuring yourself the power consumption at load, then you can start doing real efficiency comparison or apple to apple).
More importantly, as pointed out by a sibling comment, the AMD CPU in question (like most mobile Intel CPUs) has a "configurable" TDP which is set higher on most products sold. And PassMark doesn't differentiate those and only mention the "official" TDP.
To PassMark credit, they give a distribution of performance scores, just compare the distribution of the Ryzen and the M1 and you'll see (you have to scroll down a bit to see the graphs) :
Ryzen : https://www.cpubenchmark.net/cpu.php?cpu=AMD+Ryzen+7+PRO+585...
M1 : https://www.cpubenchmark.net/cpu.php?cpu=Apple+M1+8+Core+320...
In general, you can't compare TDPs even within a brand, they rarely mean what it used to mean a few years back as they "innovate" with various turbo mechanism and other OEM configurable settings.
[1] https://www.synopsys.com/dw/emllselector.php?f=TSMC&n=7&s=r3...
This is physics after all and 15W are 15W no matter if they go into an Apple M1 or an AMD Ryzen.
This is also where design decisions matter: for example, a while back I measured hashing performance for some boxes which needed to check data integrity and an Intel chip handily lost despite being faster on everything else because the embedded processor I was comparing it to had dedicate SHA hardware which was both faster and more power efficient than a generic x86 implementation. That’s ancient history now but I would expect Apple to aggressively explore opportunities to improve their stack like that since they control it at every level - for example, I believe benchmarks have shown Objective-C message passing is considerably faster on M1.
Apple has the advantage of developing and deploying hardware, OS and system software completely in-house.
AMD only supplies chips and basic firmware, both of which can be configured by OEMs/ODMs and the OS and software come from entirely different parties again.
So the usage profile depends on external factors, not just the CPU itself. In the end, however, a 15W power budget is a 15W power budget and an M1 under full load and a Ryzen under full load will have the same thermal output if configured the same (as far as power consumption goes).
How well the waste heat is managed is not in the hands of the CPU.
Finally, again, 15W TDP is not the same as 15W under normal usage. That misunderstanding appears to be driving most of your disagreements in this thread.
No. The CPUs can be configured to consume no more than 15 Watts, even if few OEMs do so. Same goes for the M1 - there's no difference with regards to this: both the MBP 13 and the Mac Mini have higher power limits and active cooling for that reason.
In fact, the latest U-series mobile Ryzen CPUs are even optimised to be most efficient at a 15W power level, contrary to Intel's Ice Lake chips, which get the most performance at a higher wattage configuration of 28 Watts.
That's a term in the industry with a specific meaning:
https://en.wikipedia.org/wiki/Thermal_design_power
The key thing to understand is that this is not measured power consumption while running the benchmarks and you cannot reliably compare the values even across the same product line, much less across chips — especially when we're talking about SoC designs where, for example, the TDP refers to the entire chip but the benchmarks being discussed are all CPU-focused and don't even exercise the GPU at all. We also know that TDP numbers are not a hard ceiling: there are some chips which under some conditions — most commonly but not always synthetic benchmarks — will exceed those figures, possibly by somewhat significant margins.
What you see there is are the power limits of the CPU (as reported by HWInfo64).
If you set the POWER LIMIT in the UEFI/BIOS, this regulates THE POWER consumption of the chip. NOTHING to do with TDP.
Why is that so hard to grasp for you? You can even measure the power rom the wall to confirm this. I am NOT talking about TDP here!
Those chips usually have 3.5-4GHz turbos, but in a fanless config, you'll never see them (and even with active cooling, you won't see them for more than a handful of seconds).
An easy way to verify it, is to measure the benchmark delta when on battery and when connected to an external power source. (M1 benchmarks remains almost the same)
Intel i7 9750H for example has a P2 of above 80W and only then can it break the 4Ghz barrier. Even though the processor is technically rated only 45W. At 45W it can just maintain the base clock i.e. 2.6Ghz on all cores.
M1 is much more efficient than any x86 chip on the market right now.
So if you change the performance settings they allow the laptop to draw 10W from the battery while plugged in for a little bit but it will throttle down to 95W to keep itself running. It still throttles which is I think the GGP’s point.
Which means, we don't really have a good way to benchmark power usage on laptops in a practical sense. We'd likely need to bust out the soldering iron + oscilloscope and measure currents entering the laptop's VRMs to accurately measure power usage over time.
I know laptops / cores have an "amp-counter" on board somewhere, but there's no guarantee that these devices are consistent or accurate across different laptops. Its sufficient for measuring how much energy different bits of code has (ex: Linux powertop tools), but not sufficient at comparing Apple M1 vs AMD Zen3 chips. We need a 3rd, trusted and independent measurement of power usage.
We can't just assume a 65W power adapter leads to 65W peak usage. Perhaps in the past when laptop designs were more in spec that was a decent assumption. But that time has passed, and today's laptops often do peak at power usages far in excess of their charger capacities (albeit temporarily, but even then, that makes measurements / benchmarks very difficult).
--------
I guess if you physically remove the battery pack (is that still allowed on these laptops?) and then plug it in, we might be getting somewhere. But the Macbook Pro doesn't have an easily removable battery pack.
That's the reason we didn't review laptop CPUs when I reviewed CPUs. You can get exact CPU power draw on a desktop motherboard (by using an amp clamp on the P8 connector) but it's hard (or not possible) to do that across multiple laptop chassis.
Removing battery (when possible) is not a solution either as what you get may differ a lot from classic "plugged in" usage (see the references to the MacBook Pro and Dell that used an i9 that still drained the battery when plugged in, because they can use more power than the power adapter brings).
On top of that, way too much depends on the OEM design and the performance of a given CPU will greatly vary from one chassis to another, because of the various throttling mechanism and the various configurable things that OEM can do (it's not just the cTDP, you can as an OEM play with various turbo times, another person mentionned P2 states, which is one of those).
So a given mobile CPU performance means nothing at the end of the day, only the laptop "as a whole" can be measured, which is why you don't see good quality benchmarks of mobile CPUs.
Anyway, just a small complement :
> laptops / cores have an "amp-counter" on board somewhere
Intel (and AMD to some measure) CPUs all have various sensors on chip that gives you the power consumption in watts (or amps, depending). They can be read with software such as hwinfo [1].
Those are usually not incredibly reliable though, they are not calibrated per CPU and it's very much a guestimate that could easily be in some cases +/- 5W off.
So sadly, not usable either (especially on mobile).
[1] : https://www.hwinfo.com/
https://www.dell.com/support/kbdoc/en-us/000140513/gaming-la...
This is 100% an Intel thing.
Similarly, benchmarking is a core component of CPU evaluation, allowing for isolated (and combined) analysis of the CPU's performance characteristics.
Handwaving both of these away is extremely misleading.
The last part is not entirely accurate. They have the same TDP. Not the same power consumption. Because 5850U doesn't use 15W in those test. The same goes to M1 which is closer to 20W Max.
The word TDP means Typical TDP by both Intel and AMD and not what it means in literal sense. That is excluding cTDP and other state like PL2.
Worth mentioning the M1 achieve those single thread performance at no more than 5W, if you put the two on equal footing, even accounting for the possible node improvement, M1 is still quite far ahead in terms of pref / watt. And the 5850 is already on Zen 3.
The next emotional preservation tactic usually cites the old GF IO Die, but that was only on H/desktop series chips anyways and furthermore they still lose to Apple in sheer performance per watt.
It's August 2021 and we still have to have this conversation. Sigh
AMD reduced the CU count down to 8 and ramped the clockspeeds which is terrible for the thermal budget. If you need to offload stuff to the GPU, both GPU clocks and CPU clocks dramatically lower. AMD needs a 16CU design with RDNA2 if they hope to actually compete with current and upcoming designs.
Speaking of upcoming, Apple's next generation will be announced in the next 2-3 weeks. A15 and either M1X or M2 (or maybe both) on N5P which should be 10-15% better than the previous N5 process. That's what 5850U is actually competing against considering how long it took to get out the door.
Things aren't looking pretty for x86. Now if we could just get some nice RISC-V designs shipping...
EDIT: not the same processor: https://simplynuc.com/cbm1r8rb/
I can't actually use them.
The M1 innovation is locked to Apple. If I have a great idea for a new device that would be enabled by the M1 perf/power/spec, I'd have no chance of building it.
I hope Google continues what they did with Coral with the Tensor chip. A Raspberry Pi style device, or a compute module of this chip would be a fever dream.
The best thing we have right now is the NXP iMX8 when it comes to performance, and still the RPi when it comes to ease of use.
Maybe if have to dig into Qualcomm and check how to get hands on their 8xx chips. The fact that there are not really SBCs with them however tells me that it's not easy to get them/use them.
I wish more articles would explicitly call this out. I've seen 100% used to mean 100% MORE more than a few times ...
https://news.ycombinator.com/item?id=28346908
On another note, I was amazed to see the SHA-1 and SHA-256 (and the Blake2b) performance numbers when running the benchmark of minio’s blake2b-simd (over M1 Golang) on a 16GB M1 Air. For some reason, Go on M1’s SHA-1 and SHA-256 are getting orders of magnitude better numbers (~2.5GB/sec!) and even the Blake2b code is beating the numbers for the same code running on an intel Air (which uses the SIMD codepath).
However the Icestorm cores also use substantially less energy so they are an efficiency win regardless. Plus they take up use significantly less physical space which is a large cost saving for the SOC part.
For my workloads it’d be an overall win to have more cores at that speed. The more the better; I’d cap out at maybe a a hundred or so.
Obviously Firestorm is better, but a hundred-core desktop CPU at present seems… unlikely.
[0] https://www.amd.com/en/products/cpu/amd-ryzen-threadripper-3...
I can’t use hyperthreading. It does give a 60% speed boost, but it’s also disabled in production so…
In the simple benchmarks, the speed differences range roughly from 2x to 5x. It looks like the current configuration (4 Firestorm + 4 Icestorm) is pretty well balanced: equisized alternatives (5 Firestorm + 0 Icestorm, or 3 Firestorm + 8 Icestorm) can be faster for specific workloads, but probably not across the board. The Apple CPU team really knows what they're doing (but that has been abundantly clear for many years now).
I wonder what this means. The efficient assembly probably has fewer instructions that use vector instruction and floating point calculations more, while the "idiomatic" Swift probably has just a larger number of instructions that aren't doing heavy calculation. Does that imply then that the high performance cores does much deeper pipelining, but the the number floating point units or whatever is probably pretty similar across both types?
Firestorm has 128KB L1 per core and 12MB shared L2.
Icestorm has 64KB L1 per core and 4MB shared L2.
At the end of the day, if all you're doing is basic adding/multiplying, you just don't need a fancy core (and hence: GPGPU, right?)
So this microbenchmark result isn't very surprising, although it's nice to see the principles in action.
Firestorm has 2x the ports across the board (int, fp, load/store, and simd). To keep those ports busy, it has a massive instruction window (630 entries). Keeping those fed requires large I-cache.
x86 instructions are 15-20% more dense than uarch64, but the fastest x86 cores still have only 32kb of I-cache (in fact, AMD went down from 64kb in Zen 1 to 32k in Zen 2/3), so I doubt doubling that is the major limitation here.
In this particular micro-benchmark, I doubt even extremely bloated code would get anywhere close to 1k of instructions for the loop let alone 64k. There might be something to say about the size of the dedicated loop cache that many chips have, but that's a completely different animal.
I spent a great deal of time, learning "idiomatic Swift" coding practices. I can do some damn clever HOF stuff (inlined maps and reduces).
This tells me that I may be better off, doing "classic C"-style coding, like I did, when I was just getting started with Swift.
Sigh...
Except with many compilers and depending on the loop, for-loops may optimise much poorer than their C++/Rust iterator based counterparts.
That's because modern language constructs can communicate intent much clearer to the compiler than a primitive for-loop can; the compiler has to make fewer assumptions (which can can be wrong).
So no, it's not "generally true" at all.
Idiomatic Swift is still great for code which doesn't have significant performance consequences and/or is likely to only ever run on the Firestorm cores—i.e. any user-interactive process.
If you're writing a background task which is expected to have some non-negligible load on the CPU, it could make sense to experiment with simpler C-style coding for the hottest loops in your code.
Icestorm cores max out at 2GHz, but have 30-50% the performance of the 3.2GHz Firestorm cores. Accounting for that frequency difference, the actual performance difference if Icestorm were to boost clocks would be closer to 50-60% of the Firestorm IPC.
Zen 3 at 4.9GHz trades blows with M1 at 3.2GHz in single-core benchmarks.
It seems when you adjust all the things that Icestorm should be 10-30% slower than Zen 3.
Putting that in perspective, Icestorm cores (based on this one microbenchmark) would have around the same performance per clock as the original Zen cores, but only consume a 5-10% of the actual power consumption.
You could do all this in other languages too, but it wouldn't be the "idiomatic" choice during development; it would be clear that you're giving up on performance for the sake of flexibility.
But for some “inner loop” stuff, it could make a big difference.
Another issue, is that visitors will hit the entire set; even when it is no longer necessary to continue. We can do some tricks to make subsequent visits shorter, but they’ll happen, nonetheless. This always bothers me.
The real problem is the existence of the overhead isn't obvious from the code.
Consider JS. The builtin implementation is naive and chaining calls results in multiple passes over the array.
Lodash is different. It uses an iterator internally, so chaining isn't much of a performance difference (just the overhead of the function calls which are probably inlined by most JITs anyway).
I see no reason why Swift couldn't (or shouldn't) do this in a future version.
There are also what appear to be intermittent sound card issues. But Everytime I make an appointment to get it fixed, it decides to not have it.
I mean it's ok, but I think wait for v2 of the chip. But I just don't get how people's experience with the M1 is so different than mine.
https://twitter.com/fwoaroof/status/1330041394019397636?s=21
With this CPU design some cores are optimised for performance (at the expense of using more power) while some cores are optimised for efficiency (using the least power at the expense of computing performance). This makes sense for laptops and smartphones, as it can save power and thus run longer when being powered by batteries. But (in my opinion) not for Desktop PC's where most people care more about computing performance than saving a few watts.
[1] https://hardforum.com/threads/killawatt-owners-whats-the-idl...
Allow package-C6. Turn on PCIe power saving stuff.
There is a bunch of stuff that mildly hurts performance and greatly improves efficiency in intermittent workloads, by letting hardware sleep/power-off when not needed.
I don't like having to run a PI for some stuff just because i don't want my huge tower running all the time, it would be really neat if it could run at anything between 5 - 600W, not sure though if the PSUs would be able to offer that range.
It's fairly clear that the Icestorm cores represent a performance gain in terms of performance per watt, but also die area. The four Icestorm cores and their support infrastructure takes up about the same physical space as one Firestorm core with its support infrastructure.
I doubt that an M1 with five Firestorm cores would perform as well as the eight cores we did get.
And from the figures in the article, it looks like 1 Firestorm core is very roughly equal to 2 Icestorm cores -
Relative to their Firestorm times, Icestorms performed more slowly by:
190% running assembly language
330% running simd (Accelerate) library functions
280% running simple Swift
550% running ‘idiomatic’ Swift
So 5 or 6 Firestorm cores could have matched the current 4 + 4 config (disregarding the valid point you made about the die area).Always. In practical use. Even for the most isolated tasks there'll be enough 'maintenance work' for the system to do, in practice, that the pool of Icestorms will justify their real estate.
If I'm mistaken, it'll be because there proves to be a maximum number of Icestorm cores for the 'maintenance work' that's called for: say if you had a 64/64 system, the idea is that maybe you get to a point where you can have 'too many cores' on the Icestorm side, wastefully so.
But that assumes they play NO role in setting up the Firestorms to perform optimally. All this stuff, all this load balancing and assigning, isn't magically done by some GrandMasterCore, it's being worked out by the same CPUs that are benefitting from the load balancing.
If you can use Icestorm cores to better set up and load balance the Firestorm cores, if you can burn unused CPU on the Icestorm side to optimize the Firestorm side… then there is NEVER a situation where you'd want the Firestorms to outnumber Icestorms. If that's happening then the Icestorms are doing some grunt work to clear a path for the Firestorms, and if you can burn more cycles to further optimize that, well…
Obviously it would be great to have 8 Firestorm cores in a desktop Mac—and this will probably be what we'll get in a future A1X—but as long as we're doing silicon fan fiction, why not ask for 32 Firestorm cores? That would be cool.
Future M1 derivatives and successors will offer higher performance at a higher price.
If you ask on those terms, agreed. But if quantified in terms of value delivered, it’s a different story. For example, if your computer can be always on with tasks like updates or if you can dedicate tasks to the slow cores there may be a value there.
Not for all workloads. Daemons and services don’t usually need much oomph, and being able to run them at double the efficiency and avoid them competing for resources with software which actually wants the higher power cores would be amazing: less context switching, less cache thrashing, less random variability, …
Hell, it would also help make low-power software more reliable e.g. foobar and the audio output don’t need much but they need it. If something else is loading the machine to the gills they can get starved and audio will start breaking down.
Or between 8 firestorm and 7 fire + 4 ice (or 6 fire + 8 ice).
In this way the CPU-intensive stuff never has to break its concentration by sullying its tiny processors with utility calculations, but keeping the observable system responsive (a much less demanding task) goes to the second, weaker computer, that's running as if nothing else is happening. So neither thing ends up being slow, because neither thing particularly has to be interleaved with the other.
Source: I'm no CPU or SoC designer, but I own a Mac M1 laptop for the purposes of compiling open source plugins to the new machines, so I've done hundreds of XCode compiles on the machine, and lots of housekeeping work in Finder. None of it seems slow, by any metric or perspective.
(And note that even Intel / AMD processors do the kind of power scaling that you described without any specialised low power cores - all cores are designed to run at varying speed. Thus, when a small task is executing, the core executing it runs it at the minimum acceptable speed and saves power. But the same core can run at full speed on demand. That's a more acceptable tradeoff for me for Desktop PCs, especially considering that ARM CPUs already are very efficient.)
The M2 processor could be 4-big 64-LITTLE, or 16-big 8-LITTLE or some other combination.
Intel don't agree. They are using a heterogeneous architecture for desktop class chips in upcoming Alder Lake processors [0]
[0] https://www.anandtech.com/show/16881/a-deep-dive-into-intels...
You're probably right about laptops and smartphones. I don't have a view on whether you're right about desktops, but I think it's probably not true for the big datacentre providers, who absolutely optimise for power and heat.
If the only niche that doesn't care about power is the desktop, and desktops are absolutely a niche when compared to the total market of smartphones, laptops and datacentres, I suspect power efficiency is coming to the desktop whether you want it or not, whether it's a positive or not, if only because of economies of scale as far as chip development and production goes.