ARM MacBook vs. Intel MacBook: A SIMD Benchmark
lemire.me
lemire.me
As I write this comment, the article's numbers are: (minify: 4.5 GB/s, validate: 5.4 GB/s). These almost exactly match my numbers under Rosetta (M1 Air, no system load):
% rm -f benchmark && make && file benchmark && ./benchmark
c++ -O3 -o benchmark benchmark.cpp simdjson.cpp -std=c++11
benchmark: Mach-O 64-bit executable arm64
minify : 1.02483 GB/s
validate: inf GB/s
% rm -f benchmark && arch -x86_64 make && file benchmark && ./benchmark
c++ -O3 -o benchmark benchmark.cpp simdjson.cpp -std=c++11
benchmark: Mach-O 64-bit executable x86_64
minify : 4.44489 GB/s
validate: 5.3981 GB/s
Maybe this article is a testament to Rosetta instead, which is churning out numbers reasonable enough you don't suspect it's running under an emulator.Update, I re-ran with the improvements from downthread (credit messe and tedd4u):
% rm -f benchmark && make && file benchmark && ./benchmark
c++ -Oz -o benchmark benchmark.cpp simdjson.cpp -std=c++11 -DSIMDJSON_IMPLEMENTATION_ARM64=1
benchmark: Mach-O 64-bit executable arm64
minify : 6.7234 GB/s
validate: 17.7723 GB/s
Note that my version also uses a nanosecond precision timer `clock_gettime_nsec_np(CLOCK_UPTIME_RAW)` because I was trying to debug the earlier broken version.That puts Intel at 1.16x and 1.07x for this specific test, not the 1.8x and 3.5x claimed in the article.
Also I took a quick glance at the generated NEON for validateUtf8 and it doesn't look very well interleaved for four execution units. I bet there's still M1 perf on the table here.
Edit: messe found the issue in sibling thread
const simdjson::implementation *impl = simdjson::active_implementation;
std::cout << "simdjson is optimized for " << impl->name() << "(" << impl->description() << ")" << std::endl;
When built for Intel/Rosetta, it prints: x86_64% ./benchmark
simdjson is optimized for westmere(Intel/AMD SSE4.2)
minify : 4.44883 GB/s
validate: 5.39216 GB/s
On arm64: arm64% ./benchmark
simdjson is optimized for fallback(Generic fallback implementation)
minify : 1.02521 GB/s
validate: inf GB/s
simdjson's mess of CPP macros isn't properly detected ARM64. By manually setting -DSIMDJSON_IMPLEMENTATION_ARM64=1 on the command line, I got the following results: arm64% c++ -O3 -DSIMDJSON_IMPLEMENTATION_ARM64=1 -o benchmark benchmark.cpp simdjson.cpp -std=c++11
arm64% ./benchmark
simdjson is optimized for arm64(ARM NEON)
minify : 6.64657 GB/s
validate: 16.3949 GB/s
EDIT: Interestingly, compiling with -Os nets a slight improvement to the validate benchmark: arm64% c++ -Os -DSIMDJSON_IMPLEMENTATION_ARM64=1 -o benchmark benchmark.cpp simdjson.cpp -std=c++11
arm64% ./benchmark
simdjson is optimized for arm64(ARM NEON)
minify : 6.649 GB/s
validate: 17.1456 GB/sLooks like -Oz bumps validate up another few percent.
% c++ -Oz -DSIMDJSON_IMPLEMENTATION_ARM64=1 -o benchmark benchmark.cpp simdjson.cpp -std=c++11
% ./benchmark
minify : 6.73381 GB/s
validate: 17.8548 GB/sIn my case I was building using Bazel. Bazel was running under Rosetta because they don't release a darwin_arm64 build yet. I didn't realize the resulting code was also built for x86-64.
I tried explicitly passing -march but the compiler rejected this, saying it was an unknown architecture. After some experimentation, it appears that when you exec clang from a Rosetta binary it puts it in a mode where it only knows how to build x86.
Pass `-arch arm64` not `-march`
You can also `clang -arch x86_64 -arch arm64` to build for both at once.
You can even go a step further and run your clang as native from bazel, via `arch -arm64 clang`.
Put it all together and you have: `arch -arm64 clang -arch x86_64 -arch arm64`.
It may seem like I'm joking but I'm not.
Why does my native application compiled on Apple Silicon sometimes build as arm64 and sometimes build as x86_64?
Why does my native arm64 application built using an x86_64 build system fail to be code signed unless I remove the previous executable?
[1] https://stackoverflow.com/questions/64830635/why-does-my-nat...
[2] https://stackoverflow.com/questions/64830671/why-does-my-nat...
The "arch" command is handy, thanks for that.
That's absolutely amazing result, and shows how wrong the current information in the article is. I hope the author sees what you did and updates his page as soon as possible.
I’m impressed that the M1 can keep up on this SIMD-optimized code, likely at much lower temperature / power use.
And even the Rosetta numbers are pretty decent.
Far inferior becomes....actually superior in many cases, even at SIMD.
It is clearly the case that the M1 CPU/SoC has a significant performance advantage in typical branchy single-core code, but much less advantage if any for certain kinds of heavily optimized numerics. Beyond that high-level summary, it’s good to dive into the details, and spark discussions.
Everyone is just now getting their hands on these chips, learning how to work with them, and trying to figure out how to best optimize for them.
> Intel/M1 ratio 1.2 0.9
> As you can see, the older Intel processor is slightly superior to the Apple M1 in the minify test.
I'd consider it as bigger news that M1 in one of the two tests chosen by the author (utf8) 10% faster than Intel, and in another (minify) only 20% slower, which is for most purposes something that most users won't even be able to notice. It's quite remarkable result. I'd surely write:
"As you can also see, in the UTF-8 validate test M1 is superior to older Intel processor, and in the minify test only 20% slower, even if Intel uses more power to calculate the result!"
-----
(Additionally I use the opportunity to thank again to u/bacon_blood who verified the initial claims and u/messe who figured out what the remaining bug in the author sources was! Great work!)
(Edit: the ratio 1.16 is from older native measurement. So I've also made an error in the previous version of this comment! I've wrongly connected that with the Rosetta 2 produced code. I've deleted that part of this message. Still the difference between 1.07 and 0.9 measured on two different setups is interesting, when another test is close enough).
Still, glad this was caught.
Why? We were not running the same config as the author. You have to supply twitter.json as an argument otherwise it uses the compiled binary itself (!) as the input due to off-by-1 errors in argc/argv parsing.
I’d be curious how the unified memory architecture shifts the cost dynamic for GPU acceleration. There’s a fair amount of SIMD work where the cost of copying to/from the GPU is greater than the savings until you get over a particular amount of data and that threshold should be different on systems like the M1.
Correct me if I'm wrong, but is this actually different from regular integrated graphics that have been in intel and amd chips for decades? I remember there being some initiatives from amd proposing similar offloading under the name HSA almost a decade ago. I don't think there are actually any software really using it.
In a recent interview (I think with the Changelog podcast) I heard an Apple engineer explain that the M1 had an advantage over previous systems in that not only did the data not need to be copied (which implies this isn't new) but also that no changes to the format of the data were needed given Apple's end to end control.
That's how I understand it works but I might be completely wrong.
> Shared Physical Memory: The host and the device share the same physical DRAM. This is different from shared virtual memory, when the host and device share the same virtual addresses, and is not the subject of this paper. The key hardware feature that enables zero copy is the fact that the CPU and GPU have shared physical memory. Shared physical and shared virtual memories are not mutually exclusive.
From:
https://software.intel.com/content/www/us/en/develop/article...
You mean ASIC I guess.
I think this was Apple's idea in first place. Instead of having general purpose computational machine why not have some general purpose alongside with specialised silicon for the most common tasks.
After all, isn't it GPU just another specialised unit? Why not have similar stuff for everything relevant?
If anything it turns out it's an argument that benchmarks mislead way too readily and first-principles arguments (which would quickly refute the idea of a 3.5x slowdown due to vector width) should always be used to double-check your working.
https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/...
[1] https://community.arm.com/developer/tools-software/hpc/b/hpc... [2] https://community.arm.com/developer/ip-products/processors/b...
It’s clear there are other workflows which some people characterize as “compute-heavy” where the M1 is superior.
Macintosh Quadra 610 DOS Compatible: Technical Specifications https://support.apple.com/kb/SP227?locale=en_US
Pictures: http://www.applefool.com/applefool/Quadra_610_%28DOS_Compati...
Blender, Gimp / Photoshop, Video Editing, LTSpice / PSpice and Matlab come to mind. These are consumer-ish workflows that benefit from linear algebra, but people want to do them on their laptops.
Hell, people are doing video editing on their PHONES these days, due to the convenience.
----------
Workstations and clusters are not affordable for the vast majority of users.
GPUs probably are affordable however. But these programs aren't really operating on GPUs yet (I mean, Blender and some Video Editing programs are... but LTSpice / Matlab are CPU-only still)
Julia, in particular, has the interesting-looking JuliaHub service in the pipeline: https://www.youtube.com/watch?v=JVUJ5Oohuhs&feature=youtu.be...
Digital Ocean moving data internally for free has nothing to do with offloading video editing from a laptop.
Also AWS and GCP are not providing a commodity service. They charge hefty margins.
Not necessarily, such math & ML inference workloads are done even done on a Raspberry pi, other ARM SBCs for numerous CV and other projects requiring edge compute.
Large data size or ML training is where I use my cluster and/or GPU computers.
AMD processors do not support AVX-512 (yet?).
What we are looking at now is some specific benchmark for a random library that who knows how makes use of things.
I would not be surprised if we see similar things added for audio DSP to support their creative software (it was alluded to in some marketing materials, haven't read much on it yet).
Other applications like language parsing and compiling is likely going to be a second class citizen moving forward, since Apple doesn't care about using their machines for general purpose development regardless of how many of us buy them for being solid Unix platforms.
Now they are way beyond that, so they can focus on what was the soul of Mac design during the System days.
Why do you say this? True, Mach is architecturally different from the BSD kernel but user space started out as NetBSD and its still fundamentally a POSIX system.
I never worked for either company but have worked with NeXT and Apple engineering teams on projects and wouldn’t say that I was working with people who took a non-Unix orientation, especially when compared, say, to Windows.
Would you not have considered AIX Unix? Or Unicos? People considered that Unix but it was more alien than the macos due to then constraints of the hardware.
Likewise, anyone doing NeXT development was focused on Objective-C frameworks all the way down to driver kit.
As Application developer on a NeXT, the tune was all about WebObjects, Renderman, EOF.
Applications like Lotus Improv, Wingz, and those being put out by Omni Group were the meat of the kind of applications that people considered to use NeXTSTEP, not BSD command line utilities.
Pretty much patent on commercials like NeXT vs Sun.
https://www.youtube.com/watch?v=UGhfB-NICzg
Or Steve Jobs opinion on UNIX users,
https://www.cake.co/conversations/rZXhqtP/that-time-i-had-st...
and later when back at Apple
https://www.computerworld.com/article/2591327/apple-hopes-to...
AIX is definitly UNIX, because it isnt' like A/UX, NeXTSTEP, OS X or iOS, where the UNIX layer is there more to bring stuff into the platform, while the main developer stack is something else.
It doesn't change the fact that there is nothing UNIX about XCode, Objective-C Kits and Swift Frameworks.
Also, macOS isn’t going to stop being UNIX based, it’s XNU/Mach based on BSD, that’s not going to change without a entirely new OS written from scratch.
What’s not UNIX about Swift/Xcode, it’s just a language / app, what don’t I get here.
Do you think they want to abandon Unix and roll their own OS from scratch?
This is the ecosystem that Apple and developers that buy into Apple ecosystem care about.
Those that were buying Macs to do GNU/Linux work were a welcomed addition in times of need, that is all.
I advise reading books like The Cult of Mac and Folklore.
I talk about the culture of the application developers and what Apple developers that are on the platform since the System days care about, and keep being told Mac OS X is an UNIX.
Of course it is one, that is not the point being made.
I am talking about the Apple developers culture, from those developers that care about Apple platforms, regardless of what powers the bottom layer of the OS.
UNIX can exist until the end of days at Apple, that is not what matters to Apple application developers.
As for GNOME and KDE, they are lego pieces on Linux, a fragmented experience where the command line is worshipped, for most users running something else doesn't matter, or they even change environment every couple of days, this is not what Apple culture is about.
You're right from the point of view of a GUI application developper, the UNIX core is somewhat hidden under intermediate layers, but then it's also the case for applications on Linux. Using GTK or QT, you don't deal with low-level kernel APIS much beyond POSIX either. And you can also do that on macOS, so it's a bit pointless as a purity test.
It seems difficult to argue with a straight face that Apple's developers working on the kernel, low-level layers and system libraries don't care about UNIX: that is their whole job. And as a user, you can have UNIX and decent GUIs.
macOS being UNIX was never in discussion, as mentioned it helps sales.
That's a bold statement that I've heard repeated over many years and never actually seen any real evidence of considering the Mac is the only place you can develop applications for the majority of Apple devices. Not a week goes by on HN without someone saying "Apple doesn't care about developers and/or will stop making general-purpose Macs" and yet there isn't a XCode for iOS or Android or Windows or Linux.
You pretty much don't develop anything else on Mac OS. No web, no embedded, no Linux, no Windows, no nothing. Only Mac OS software, iPhone apps, etc. I am exaggerating slightly, but you get the idea.
Or when you do, you use your Mac OS laptop as a terminal. Or at best, something to run VMs, and in both cases you don't develop with Apple OS / tools, it's just hosting or giving you non-integrated access to completely different systems.
Apple things are a pretty much distinct and closed ecosystem, and one which is quite limited to (some) endpoints.
Apple has never stopped me from installing VSCode or Atom or Sublime or Jetbrains and they've never given me any reason to think they would do so in the future, especially given Rosetta 2. I've never run a VM on my Mac.
I have no idea what point you're trying to make here, so I'm trying my hardest to argue against the best possible interpretation, and that interpretation is still so unbelievably wrong that I still feel like I must be missing the point.
So I already wrote that I was exaggerating slightly, and I still believe your case fall into that. For example about embedded dev, I had in mind things less individual tinkerer oriented and more productized. I don't know: set-top-boxes, base stations, software for trains, software for washing machines, smartcards, etc; or even big equipment controlled by an OTS desktop/laptop-like computer, or a PLC. I'm sure in a few exotic cases Mac will be involve here and there, but lets be honest, Windows as an host dev station is far more probable. Maybe a few Linux too, but probably far from the majority. And Mac OS would be probably: very far.
Now about running Intel code, I know about the excellent x86 emulation layer Mac OS+M1 have, and it's great, but I actually don't care at all for what I was thinking about, I think that broadly apply for x86 Mac as well. I'm more thinking about the software ecosystem, the precise HW CPU for devs is only interesting maybe for people developing SIMD code or, well, running VMs.
About running VSCode & co, that's great to but where are the toolchains for Mac OS host for the targets I talked about? That's why I qualified Mac OS in this case as merely used as a "terminal", I was speaking in the broad meaning of the term, a graphical terminal, not just a VT100 like terminal. The actual toolchains are elsewhere.
About web-dev, I admit that's probably where you can do most of the non-Apple-only dev while staying really native, although probably not if you need a complex server-side setup. Arguably I went way too far when I wrote "no web".
Well, nothing is absolute, and I know Mac OS remains a general purpose OS even able to host some serious dev. Just I think it is not really the most used one outside of let's say client related consumer tech and some pro-desktop tasks, mainly on Apple techs. Claiming embedded in the general case would really be stretching the narrative.
You are exaggerating wildly, and I am not sure what your idea is. Why is web or embedded dev impossible on a Mac?
But embedded: where are the most used tools for FPGA? where are the compilers for microcontrollers? Of course you can do some, in a limited capacity, with a subset of targets and a subset of toolchains, not the most used on earth. But broadly in that area, yeah you don't use Mac Os.
XCode is a good example of what I mean. It's a terrible developer experience for anything but targeting Apple's platforms in the way you want to (meaning writing Swift, using their frameworks, and their devtools, and keeping your code base small to keep their devtools functional, and never caring about that code running on a different OS).
Compare to Visual Studio, which also only runs on Windows, yet is a pretty damn good developer experience and not a nightmare to support for cross platform projects of late.
It was and is built as an IDE for Apple's platforms/languages, so this point is moot, and doesn't prove anything general about macOS as a platform.
It's like saying Emacs is bad for developers, because SLIME doesn't really play well with Javascript coding...
Again, this is repeated here at least once a week for the how many years I've been visiting this site and it's no closer to being true now than it was back then. If it was even remotely true Apple would have left the Macs on Intel and just phased them out in favor of the iPad Pro, but instead they spent billions making their iPad chips run MacOS and desktop applications and code compiled for Intel processors and real talk here, what about that gives anyone any indication that they're planning on throwing all of that away?
They're actively doing the opposite of what you're claiming and spending billions of dollars to do it, and one blog post that says "this isn't even a big deal" is all it takes to convince you otherwise?
This goes deeper than developer tools. The documentation for their core frameworks have been purged from official sources and is relegated to deprecated sites and comments header files tucked away in /Library. Kernel modules are being deprecated. There's no alternative to IPP or MKL on ARM for the M1 chip. Docs for plugins architectures are being more and more hidden away, and Apple's Developer Conference consistently only focuses on consumer facing applications - while support for professional applications and advanced computing is only available if you work for a partner organization, making it less accessible.
The reason I say that Apple is making their platform harder to use for general purpose computing is from my experience shipping code on MacOS for the last decade. It's fine if you disagree, but that hasn't been my experience. Every year it costs me more time and money to target Macs than the year before.
Custom Silicon -> Programmable Logic --------------> Software.
You see that in RF design where tech is developed with software defined radios, then moved to FPGA's/Gate Arrays, and then custom silicon. And the power requirements drop by 95% each step.
Also microprocessor manufacturers have always pushed back against specialized coprocessors as much as they can. Apple bringing it all in house allows them to nullify that impediment.
Ah, the endless Apple-hating complaining...
Benchmark here: https://github.com/bwasti/mac_benchmark
If anyone has the new 2020 MBP with AVX512 I'll update the benchmark to include avx512 instructions! I'm very curious about those numbers
Idly speculating, if this architecture favours characteristics of higher-level languages (I'm thinking of the widely-reported measurements of primitives used in automatic memory management), relatively disfavours straight-line branchless SIMD-ready streaming algorithms, and is also shipped with a matrix-friendly neural coprocessor... could that invite a change in the types of programs that perform the best?
That is, is it possible that good algorithms nicely structured and straightforwardly written in higher-level languages with good separation of concerns might actually get the most benefit? Or is this a daydream?
My educated guess is that the latter is most frequently used as a technique.
Also traps are caused by the very instruction being executed so you cannot complete the current instruction in all circumstances before vectoring to the trap/interrupt handler.
Apple nailed this, unbelievably. I have an M1 MacBook and it’s truly black magic. It’s already the fastest computer I’ve ever used. I’d rather have stability during the Rosetta 2 stage than added complexity and bugs by introducing a new bleeding edge instruction set.
The other side of this is that we're now in a window where performance-sensitive developers are motivated to start looking at NEON and adding NEON support to existing SIMD code. Which is fine... but if SVE is the long-term answer and SVE will (as its goal) support better performance scaling of SIMD code with future processors, it seems like there's a real motivation for Apple to push developers into doing things "the right way" as early as possible.
Assuming (!) that SVE is indeed the future, missing it for the first few generations is shades of 32-bit-Intel-for-Apple -- a stop-gap with some remarkably long-tailed support costs, compared to jumping straight to x64.
It's the kind of fit & finish that I would expect Apple to obsess over.
Anyway there’s already AMX
You can read what Google did on Android phones to encrypt drives here for instance: https://lwn.net/Articles/776721/
For esoteric/custom crypto it could play a part though but you have to have good reasons to not want to use standard crypto at higher speed for it to be your use case which is why I say it'd be uncommon.
Modern designs tend to be very SIMD-aware, such as BLAKE3[0] and Gimli[1].
[0]: https://github.com/BLAKE3-team/BLAKE3/blob/master/c/README.m...
[1]: § 5.6 https://gimli.cr.yp.to/gimli-20170627.pdf
(And PyTorch can be and often is run on pure CPU.)
You have to understand that current software exists because it solved a goal. If you have to first have your algorithm implimented in hardware then nobody can make anything. Without SIMD these projects simply wouldn't be able to exist. That's the hard math of it.
Discounting their existence is entirely unfair, as one of the whole points of Apple Silicon is to give Apple the opportunity to put whatever hardware into their computer that accelerates the use cases they envision for their computers. Dedicated hardware is way more power efficient that software implementations.
However, what happens if you work in a video codec that Apple didn't build in hardware support for? Software video codecs depend heavily on SIMD instructions to be performant.
FWIW, NVIDIA has made significant improvements to quality for their hardware encoders in each of their last 3 generations, and you definitely saw reviewers and creatives talking about that in particular when it came to purchasing decisions.
Apple's encoder is probably quite good at least, but I don't think it's meaningful to consider it for most benchmarks. The scenarios where you both are willing to use the hardware encoder and care about how fast it is are relatively few and far between - if you're just doing a zoom call all that matters is whether it can pump out 60fps and how good it looks, not whether it uses 3% cpu instead of 5%. I'd rather see quality/bitrate comparisons of their encoder with x264, not benchmarks.
The tradeoff with using it is basically between instruction density, power (AVX-512), and latency (GPUs are seriously powerful, but getting the data going takes time and a lot of driver bullying).
These instructions can be beneficial when leveraged providing significant speed ups for larger transfers.
ripgrep does the same, except for Intel at least, its memchr is implemented in Rust using SIMD intrinsics explicitly. And it also has a specialized SIMD algorithm (taken from Hyperscan mostly) for dealing with multiple patterns: https://github.com/BurntSushi/aho-corasick/tree/8b479a60906d...
Hyperscan takes this to a different level though. It has oodles more SIMD. I should have mentioned it in my original comment.
I was curious mostly because I never recalled any SIMD intrinsics in GNU code (ok, probably GIMP has them, so maybe I should say GNU utilities), so that would be a first.
It's interesting how much stuff leans on memchr, shame there aren't systematic wider versions taking more bytes to avoid false positives for longer literals (ignoring wmemchr): these could be nice and fast with SIMD.
And then of course there's PCMPESTRI (and its variants), but that has largely been a failure because of its high latency. :-( That's a shame, because that instruction does accept substrings up to 16 bytes.
Intel: 2x256 = 512 M1: 4x128 = 512
I'd expect Zen3 would beat them both pretty easily given that it has 4x256, as would any of the Intel chips with 2x512.
Intel better be widening its execution engine, they have focused on MgHz for far too long..
>In some respect, the Apple M1 chip is far inferior to my older Intel processor. The Intel processor has nifty 256-bit SIMD instructions. The Apple chip has nothing of the sort as part of its main CPU. So I could easily come up with examples that make the M1 look bad.
>I am not saying that the Apple M1 is not great. It is.
It's sad that this needs to be written.
[0] https://lemire.me/blog/2019/07/10/parsing-json-using-simd-in...
Running things on some accelerator (gpu, etc.) usually involves writing a specific kernel in a language subset, manually copying data and generally long latencies. Unless there is a lot of data it won't be faster.
With AVX in the best case the compiler can just vectorize some loop, speeding it up 5x without any added latency or source code changes.
See libjpeg-turbo, ffmpeg, crypto, hashing in general, ripgrep, simdjson, ...
https://woboq.com/blog/utf-8-processing-using-simd.html
See also: https://www.reddit.com/r/compsci/comments/4cq0ls/when_is_sim...
If simd was obsoleted by GPUs intel & amd wouldn't keep introducing wider simd extensions
Curious to think about how unified memory may change the ratio of flops/memory access when it makes sense to shift job from CPU (better for low number) to GPU (better for high ratio)
The Pinebook Pro is really good. Not excellent, but really good.
And seeing as I prefer FOSS, and general computing over walled gardens, it's superior to the Apple alternative.
And because Apple users have given the keys to their hardware to someone else, they get things like this happening:
https://techcrunch.com/2020/11/12/macos-apps-wont-launch/
It will only get worse; Apple has little incentive to support general computing over a walled garden where they get a cut of every software purchase.
About 92% of all iOS apps on the App Store are free; pretty sure the numbers for macOS are similar:
https://www.statista.com/statistics/263797/number-of-applica...
I use mostly free/open source software for web development on macOS; the outage didn't affect me at all.
Details on what happened: https://eclecticlight.co/2020/11/16/checks-on-executable-cod...
My mac-using coworkers were affected and lost most of a day's productivity. They'll lose more in the future.
Almost every developer I know will stop using macs if macs start actually preventing them from running whatever software they want.
The few doors that remain open will close; Mac will move to a closed market before long unless regulators prevent it, or that OS is retired entirely.
The more closed it gets, the more customers they will lose. Right now macs aren't closed at all, so honestly the customers they've lost so far (yourself included) are... faint-of-heart? Excessively sensitive? Not sure what the best way to phrase this is. You've effectively stopped buying macs for something that could happen, not something that has actually happened. I would find it incredibly surprising if you install enough unsigned GUI apps for that whole warning + have-to-open-it-with-right-clicking thing to be a dealbreaker for you on its own. And if it was, they could've lost you at any time because that seems like incredibly fickle consumer behavior.
An actual closing of the platform though will be a watershed moment.
Apple is well aware that most of its recent customer growth, and almost all its future customer growth, is in iOS and similar walled gardens.
The Mac users are proving to happily accept greater restrictions on ability so long as their preferred tools continue to work. You state as much yourself: running unsigned software is surprising to you; as though that should be considered abnormal behaviour. Consumers of Mac products will happily accept greater restrictions if they can be convinced it brings quality; whether or not it is successful in doing so.
Again, and for the final time, there has not been any restriction on ability. You can still run the things you've always been able to run. You just have to go through a warning first. That is not a restriction on ability. You know Windows does similar now, right?
Running unsigned software is not surprising to me at all. Running so much unsigned GUI software that it becomes restrictively annoying... is absolutely surprisingly. I guarantee you I run a lot more unsigned GUI apps than the vast majority of users, and it's still such a minor inconvenience I barely notice.
It's not even a particularly bold prediction; it is completely in-line with their whole product tragectory.
You will accept OSX turning into iOS, and call it innovative, powerful and unburdened; because that's what Apple will call it.
I wonder what final cut and other software suites tailored for the new Macs are doing under the covers as they seem to perform very quickly. Also this isn’t really all the relevant as cryptographic validation is required more and more on the server side of things whereas Apple doesn’t have a SKU for the server market.
On iOS via the Accelerate framework.
I doubt many would bother dealing with NEON intrics directly unless in some performance critical cases, though.
https://developer.arm.com/architectures/instruction-sets/sim...
So x86 is certainly not going to adopt it.
The GPU/neural stuff is basically nothing else but SIMD, just in slightly different guise and running on different hardware.
On PC it is the difference between MMX/AVX (CPU) and shaders/CUDA/"compute" (GPU), on M1 which is all-in-one SoC it is just a different part of the CPU being used, optimized for somewhat different tasks.
SIMD is only a technical term for one type of parallel computation - single instruction, multiple data. It is not some sort of Intel-specific magic technology (and the term far predates Intel's support for it, going back to vector processing on Crays and such).