And that's the big stuff! There's also the steady incremental improvements such as battery technology, SSD and RAM that are ongoing.
You also have an incumbent (Intel) who's lost their way. If they were as scrappy as they were 15 years ago this wouldn't feel like such a banner year.
Yeah, I think we are in a new era of computing.
This actually reminded me of how I got skeptical comments from people when I told them that my small and scrappy Pentium-M (P-M -> M1, HAH!) based laptop was almost as fast as their desktop P4 monsters in compiling code.
Northwood was a relative bargain when you cranked the bus clocks up well beyond what it said on the box.
That's great for shareholders.
Meanwhile, keep in mind that IBM exited the consumer computing device market. If that's also in store for Intel then it's a bit pointless in the PoV of those in the market for consumer computing devices.
That is called Lindy effect: https://en.wikipedia.org/wiki/Lindy_effect
I'm not someone into the inner-working of chips much, but is "ray tracing" a new term used for something in microprocessors now? Or is this the same graphic "ray tracing" we were doing back in the 80's on Amigas and Atari STs?
This is still not perfect and many ray tracing techniques rely on accumulation over time which limits certain images from working (I imagine raytracing a small particle cloud, or fast shifting objects to be a worst case scenarios)
Yes, the NVIDIA 2xxx RTX series had it two years ago, but this is the year where it's actually viable and not so gimmicky.
But when every PS5 and XSX has raytracing hardware, suddenly it makes sense. That's going to be helpful for getting it supported in PC titles sooner.
I remember playing with a number of demos on my intel core 2 duo macbook (not pro) a decade ago.
I've read stuff about how the M1-packing MacBook Air shipped with a SSD whose burst write speed was far higher than the one shipped in older MacBooks. The bottleneck on build jobs tends to lie on disk access, specially with projects comprised of a significant number of small files whose build also outputs a bunch of small files.
This is one of the reasons behind doing builds on RAM drives.
If that's the case then these weird speedups might not be due to magic properties on Apple's M1 professor but due to the fact that the processor doesn't idle as much while waiting for all those reads and writes to finish.
If anyone has any data on this, please do share.
From the TechCrunch review, which is pretty breathless but also contains a lot of good data: https://techcrunch.com/2020/11/17/yeah-apples-m1-macbook-pro...
There could be all manner of processor optimisations that are taking time to process. Like for like would be much more indicative
I think that for consumer products, which these are targeting, software responsiveness, usability, and battery lifetime are by far the most important metrics.
These chips help with battery, and can hell with software responsiveness if there's is developer focus on it.
But what will probably happen is that development teams will buy the fastest computers they can, and then develop software that is mostly, somewhat adequately performing on this beefy hardware, then ship it to customers that are on weaker hardware.
This effect is especially pronounced for web software, where it's easier to make unresponsive interactions due to so many layers of software, especially with developer network connections usually being 10x-100x less latency than users.
The 8GB Macbook Air at $899 educational is faster, and will feel faster, than any laptop that anyone has owned or thought of owning at that price point.
Millions of buyers who need laptops for *-at-home activities will sing its performance praises on Sheets and Salesforce.
The "I need 16GB crowd" of content creators need more RAM and GPU but they are a tiny fraction of the market for laptops.
MS Office apps, for example, are horrifically unresponsive on Macs. Switching the ribbon to a new view has 700-1000ms of lag on my 2.4 GHz i5. Maybe an M1 brings it to 350ms. Once MS developers start developing on an M1 laptop, the developers will change code, and it will slow down, and until it gets slower than it currently is, the code will not be optimized.
This is what I mean about software being like a gas rather than a liquid. Any new CPU performance will be consumed by developers because their threshold for performance optimization changes with each new performance improvement.
If they were ever going to be fast, they would already be fast. They are a software problem unto their own.
I think there are two routes to making software faster for users: 1) intense education of developers and rewarding them properly for keeping software responsive, and 2) only letting them develop on 5-10 year old hardware. I'm not really sure 1) would work with many teams, but I'm pretty sure 2) would.
The average Joe user just uses a browser and something like Spotify. Even most word processing by college students is in Google Docs now - very few people I knew bought MS Office for their Macs when I was in college 5 years ago, even with a $99 student license through the school.
Developers will use all available resources until their is pressure to be more efficient. This is not a critique of developers, this is the nature of software. Unless critical development time is spent making sure that software is responsive, it will only ever have barely acceptable performance.
Which is why new, faster CPUs have very little effect on users. Any performance gains will be gobbled up by new software frameworks that promise better use of developer time, but which may come at an absolutely tremendous cost of UI responsiveness.
Spotify, Slack, Office on Mac, hyper complex JavaScript web frameworks... all will continue to take more and more CPU cycles that are available.
Apple is just jumping onto their existing ARM track, once they migrate their product line, which has surpassed Intel. Once they've migrated all their lines to ARM, the performance gains will be more like they have been on the iPhone/ iPad over the past few years. Mostly 20-30%/ year.
Though I suspect if Qualcomm were able to source a TSMC 5nm chip, it would be more competitive with Apple than Intel is at this point though. Apple has a lot of other things going for it where Qualcomm lags (the Secure Enclave, graphics performance, audio and photo processing, the neural engine etc etc)
Intel gets no such benefit of the doubt. I have no idea what on earth is going on over there.
What I was trying to get at is that the ARM designs plus the TSMC fabs are a big part of Apple's success here. The pieces are out there where someone else could put together an ARM based package that's more competitive with Apple. In retrospect, maybe it's more likely to see something like this from Nvidia than Qualcomm.
Even then, it's hard to say how competitive that CPU would be. Just based on Microsoft's Surface with it's half-assed Qualcomm CPU, it seems feasible though.
I'm guessing the high performance ARM chips for non-Apple devices will be coming from Nvidia or Samsung in the future
What does matter, IMO:
- assembling a killer team
- 5nm process
- high speed, low latency DRAM
- big-little
ARM has been improving much faster than Intel.
Apple has been executing ARM much better than anyone else.
Apple's auxiliary processors and integration have been top notch.
TSMC has been crushing Intel in getting to 5nm.
If these are actually the reasons for the performance difference, and it's difficult to do these on x86 because of the instruction set, it seems to this amateur that ARM64 really does have an advantage over x86.
"A weak memory ordering model, like the one in Apple silicon, gives the processor more flexibility to reorder memory instructions and improve performance, but doesn’t add implicit memory barriers."
[1] https://developer.apple.com/documentation/apple_silicon/addr...
Here's a kernel extension someone built to manipulate this feature: https://github.com/saagarjha/TSOEnabler
> Other contemporary designs such as AMD’s Zen(1 through 3) and Intel’s µarch’s, x86 CPUs today still only feature a 4-wide decoder designs (Intel is 1+4) that is seemingly limited from going wider at this point in time due to the ISA’s inherent variable instruction length nature, making designing decoders that are able to deal with aspect of the architecture more difficult compared to the ARM ISA’s fixed-length instructions.
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
While this is wonderful for ARM in the now-term, we just moved from walled ISAs to a plurality of ISAs, compute just became a bulk commodity in a way that it could not with an x86 duopoly.
Anyone can now take off the shelf RISC-V designs that are currently at > 7.1 coremarks/mhz and get them fabbed on Glofo or TSMC. If you need integrator help, you can use the design services of SiFive.
Back when the iPad Pro with the A10X came out, Apple claimed it was faster than half of all Laptops sold and people in the PC space were yamming on and on about how numbers don't show how much better x86 cpus are at 'desktop stuff' and that ARM cpus can't equal x86, even with the same thermal envelope and shouldn't ever be compared. Ironically, many are now stating that the reason why they are so good is because of ARM, which isn't true either lol.
It needs 30W at 4 cores 3.2Ghz. Ryzen needs around 5W per core but it's on a worse process. The entire system does use less power than a x86 system but that has nothing to do with the processor. It's more about how the SoC is arranged and that RAM is (almost) on the same package. It means they can get away with higher bandwidth and lower power consumption for the entire system.
The idea that it's all about the processor is completely wrong. Yet all we have heard is how fanboys cry it's going to be 3x faster than desktop CPUs because of misleading TDP numbers.
I'm not sure this test is deserving of the breathless headline and commentary, especially since the original tweeter later follows up with:
> Extra info: The M1 macbooks (air/pro) can't drive 2 external screens, and the air throttles a bit after 3+ minutes sustained compute (20-30%)
https://twitter.com/rikarends/status/1328753176552632321
I'll be more interested if the M1 can compile something 2x quicker than the i9 when the compile time on the M1 exceeds 30-60 minutes rather than being less than a minute.
EDIT: to be clear, I expect the M1 to feel faster than the i9 for the vast majority of users, however the headline is "in a real world Rust compile", implying that this is a more valid test than synthetic benchmarks. I take issue with that, as I don't really consider something that compiles in less than 2 minutes on 6 year old hardware to be a much more useful test than the benchmarks.
We already know the A-series of chips performs incredibly in short workloads. We have no information yet on how it performs under sustained workloads.
What makes you think that given sufficient cooling, it will not perform exactly the same as the M1 in the MBA but sustained? It’s not like the ARM architecture changes anything in the thermodynamics of cooling cpus compared to an x86 chip, right?
I’d wager that under load an i9 with passive cooling wouldn’t even last 30 seconds without throttling below even its base clock, if it doesn’t just shut down to prevent frying itself
Asking about how it would do in a computer with sufficient cooling is about as relevant as asking how it would do in a computer with a usable keyboard or OS.
So again, what would make anyone think that an M1 with decent cooling would not be able to maintain the current ST performance indefintely, or a hypothetical 8+8 or even 16+16 core M1X or M2 with a TDP of 100W and top-notch cooling solution would be impossible?
Don't forget that scaling up is also not just about frequencies, there are also packaging considerations - the CPU dies have to actually be able to dissipate the heat generated, and the package itself has to be able to do so as well. I'd expect that this is something AMD and especially Intel have a leg up over Apple with - although, considering they've already tread that ground it makes Apple's job a bit easier too.
https://techcrunch.com/wp-content/uploads/2020/11/webkit-com...
https://blogs.gnome.org/mcatanzaro/2018/02/17/on-compiling-w...
What's also hugely impressive is that under better cooling conditions, it's also 25% faster.
None of these numbers capture headlines like 2x sadly, but that's still massively impressive.
Curiosity got the best of me too and I ran the test on my late 2013 MBP, 2.3 GHz i7, 16 GB ram. Compilation took 44 seconds and the fans didn't even spin up (with 23 ºC ambient temperature).
A little further down the thread [0] he gives the actual numbers, which are around 20s on the M1, which puts the i7 [1] at around 40s.
I'm not sure how much of this is Rust specific, but for my own projects I haven't noticed a big difference between my mbp, an old i7-3930k and a newer i5-8500. The MBP is somewhat slower, but it only has 4 cores while the others have 6.
[0] https://twitter.com/rikarends/status/1328706132752347138?s=2...
[1] A tweet corrects the MBP CPU as being an i7 and not an i9.
Which was Tesla's equivalent of Jobs walking on stage with the iPhone, itself an homage to his iMac G3 turnaround, in turn a recapitulation of his promotion of the Apple II.
Even their most inexpensive products are on the higher end of the pricing spectrum, and they could be 5x the performance but it still wouldn't matter.
There's a reason why Chromebooks are so popular and it sure isn't anything to do with their performance.
Nope, LPDDR4x-4266 is LPDDR4x-4266. Apple, Intel, and AMD all have access to the same RAM. The Firestorm core is the real advantage.
https://www.anandtech.com/show/16252/mac-mini-apple-m1-teste...
“One aspect we’ve never really had the opportunity to test is exactly how good Apple’s cores are in terms of memory bandwidth. Inside of the M1, the results are ground-breaking: A single Firestorm achieves memory reads up to around 58GB/s, with memory writes coming in at 33-36GB/s. Most importantly, memory copies land in at 60 to 62GB/s depending if you’re using scalar or vector instructions. The fact that a single Firestorm core can almost saturate the memory controllers is astounding and something we’ve never seen in a design before.”
And I'm really hopping Intel manages to get its act together. Competition is great for everyone.
But thanks, I will :)
If you want a more realistic sense of how the M1 performs relative to x86 peers in raw, equivalent workloads, there are some cinebench numbers appearing out there:
https://hexus.net/tech/news/cpu/146878-apple-m1-cinebench-r2...
I am way more interested in non-accelerated, deeply out-of-order instruction processing capabilities if we are talking about a "new era of computing". Being able to compile to ARM faster than x86_64 and having super fast HW codecs for processing special unicorn byte streams are not very compelling arguments in this context.
Show me an ARM chip scaled up to the same power+price budget as an Epyc 7742, throw them both at a 24 hour SAP benchmark, and then I will start paying attention if the numbers get close.
Not saying it isn't a novel form of transport or equally useful in most cases, but... we're comparing a very stripped-down SoC to systems that have vastly more complexity for several different reasons - not least the ability to support more modular CPU/memory/GPU configurations.
Rework the existing x86 cores that we have into similar configurations and we'll likely see pretty substantial efficiency gains there too.
In IT though, I don't expect any changes over to ARM hardware for another decade. I know where I work we have plenty of legacy cruft which would probably run into some weird edge cases if emulated on ARM.
I think the bigger question is what does this spell for x86-64?
I also have my doubts, but would be great for the market. (New Exynos 2100 supposedly also being up there.)
So Qualcomm won't be quite so far behind Apple, but it's still pretty significant.
If we'll just look at geekbench's single core bench, I'm sure Apple will still lead. (And overall likely still produce the better chips.)
Not much?
People are acting like M1 destroys x86, but as AnandTech showed in the recent benchmark, M1 is trading blows with Zen 3 in single thread performance while having much larger core and process advantage (5nm vs. 7nm) thus being actually more expensive to produce.
Are there any workloads that requires or perform better with x86?
The difference might be way smaller or non-existent once we see some comparisons with latest AMD offerings (especially when we have also CPUs with comparable process).
The M1 isn't a tiny power-sipping mobile part - it sucks power down just like AMD's 15W TDP CPUs do. The efficiency gains Apple are getting here are likely to be the result of several factors, only one of which are the CPU cores themselves.
Single core performance does not differ much between mobile and desktop CPUs these days.
You also can't compare manufacturing cost and retail cost. But given AMD cores are both smaller and manufactured on older process, it's probably safe to assume AMD ones are cheaper to manufacture.
“But this benchmark has these issues”, “but this is hardware accelerated, so it's not a fair comparison”, etc. I think when enough benchmarks and real world usage have been run, it will sink in.
What did people expect? They've been killing it with 5W fan-less chips for years. Have you seen how confident Apple is in those videos?