Apple's M4 has reportedly adopted the ARMv9 architecture
wccftech.com
wccftech.com
I wonder if there is licensing at play though. Apple may have gotten a really great licensing deal on ARMv8 that they wouldn't be offered for ARMv9.
https://www.electronicsweekly.com/news/business/finance/arm-...
However, Apple basically commissioned ARMv8 in the first place to develop the A/M chips, so that presumably helps.
They used the ARM610 in the Newton in 1993 (ARM was founded in late 1990) and then an 8 year gap to the iPod in 2001 (ARM7TDMI which are ARM designs.) Their first in-house ARM design (I believe) is the iPhone 4 in 2010.
They definitely didn't "architect/design ARM" for nearly a couple of decades after founding ARM, yeah, but they did use them.
Yes. Watching it unfold, the whole misinformation self spread, and back into its loop is both interesting and tiring. It took months and some hard work to camp down Unified Memory being SRAM and something special. But this ARM deal? 5 years and counting.
It is two point;
Apple has an architectural license, which somehow became a "special deal" as if they are the only one doing it. They are not and even the Amphere Computing has one.
Apple was the part of the founder of ARM, and somehow this gave people impression they have a "special deal".
And had hajile not step up and provided the ARMv9 being a superset of ARMv8 etc etc I would have to spend some time to look up the detail of superset just to stamp out these type of non-sense. ( Each ARMv8+ and v9 have way too many features and optional extensions I dont even remember which is which )
And this is not the first time YouTuber Vadim Yuryev gave something out that is completely wrong.
Edit: Only Apple has implemented SSVE and SME I think.
And implementing 256b SVE has different issues depending on how you do it. 4x256b vector ALUs are more power hungry than generally useful. 2x256b is only beneficial over 4x128b if you're limited by decode width, which isn't an issue now that A32/T32 support has been dropped. 3x256b would probably imply 3x128b which would regress existing NEON code. And little cores don't really want to double the transistors spent on vector code, but you can't have a different vector length than the big cores...
If you have four 128-bit packed SIMD, you must execute 4 different instructions at once or the others go to waste. With SVE, you could (in theory) use all 4 as a single, very wide vector for common operations if there weren't a lot of instructions competing for execution ports. You could even dynamically allocate them based on expected vector size or amount of vector instructions coming down the pipeline.
Additionally, adding two 2048-bit vectors using NEON (128-bit packed SIMD) would require 16 add instructions while SVE would require just one. That's a massive code size reduction which matters for I-cache and the frontend throughput.
Modern out-of-order cores are already good at superscalar execution, so why not let them do their job? 4x128b units give you much more flexibility and better execution granularity.
That aside, see the "cmp" sibling thread for a major (4x penalty) downside to 4x128.
Could you point me to the "cmp" thread you mentioned? I don't know where to look for it.
Sure, it's https://news.ycombinator.com/item?id=40465090.
I'd already classify this as "very wide". And the story is far from being that simple. Intel's 512-bit implementation is very area- and power-hungry, so much so that Intel is dropping the 512-bit SIMD altogether. AMD has 4x add units, but only two are capable of multiplication. So if your code mostly does FP addition, you get good performance. If your workflows are more complex, not so much.
The thing is that on many real-world SIMD workloads, Apple's 4x128bit either matches or outperforms either Intel's or AMD's implementation. And that on a core that runs lower clock and has less L1D bandwidth. Flexibility and symmetric ALU capabilities seems to be the king here.
> Sure, it's https://news.ycombinator.com/item?id=40465090
Ah, that is what you meant. Thank you for linking the post! My comment would be that this is not about 128b or 256b SIMD per se but about implementation details. There is nothing stopping ARM from designing a core with more mask write ports. Apparently, they felt this was not worth the cost. Other vendors might feel differently. I'd say this is similar to AMD shipping only two FMA units instead of four. Other vendors might feel differently.
AFAIK it has not been publicly disclosed why Intel did not get AVX-512 into their e-cores, and I heard surprise and anger over this decision. AMD's version of them (Zen4c) are a proof that it is achievable.
I am personally happy with the performance of AMD Genoa e.g. for Gemma.cpp; f32 multipliers are not a bottleneck.
> The thing is that on many real-world SIMD workloads, Apple's 4x128bit either matches or outperforms either Intel's or AMD's implementation
Perhaps, though on VQSort it was more like 50% the performance. And if so, it's more likely due to the astonishingly anemic memory BW on current x86 servers. Bolting on more cores for ever more imbalanced systems does not sound like progress to me, except for poorly optimized, branch-heavy code.
I looked at the paper and my interpretation is that the performance delta between M1 (Neon) and the Xeon (AVX2) can be fully explained by the difference in clock (3.7 vs 3.3 Ghz) and the difference in L1D bandwidth (48byes/cycle vs. 128bytes/cycle). I don't see any evidence here that narrow SIMD is less efficient.
The AVX-512 is much faster, but that is because it has hardware features (most importantly, compact) that are central to the algorithm. On AVX2 and Neon these are emulated with slower sequences.
Isn't the L1d bandwidth tied to the SIMD width, i.e., it would be unachievable on Skylake if also only using 128-bit vectors there?
That is interesting! So do I understand you correctly that the 512b vectors allow you to implement the algorithm more efficiently? That would indeed be a nice argument for longer SIMD
> Isn't the L1d bandwidth tied to the SIMD width, i.e., it would be unachievable on Skylake if also only using 128-bit vectors there?
It's a hardware detail. Intel does tie it to SIMD width, but it doesn't have to be the case. For example, Apple has 4x128b units but can only load up to 48 bytes (I am not sure about the granularity of the loads) per cycle.
I agree that the number of L1 load ports (or issue width) is also a parameter: that times the SIMD width gives us the bandwidth. It will be interesting to see what AMD Zen5 brings to the table here.
Can you explain why thats bad?
Don't you still get full utilisation of the 4x128b units?
This is why I like the ARM/Apple design with "regular SIMD" and "streaming SIMD". The regular SIMD is latency-optimized and offers versatile functionality for more flexible data swizzling, while the streaming SIMD uses long vectors and is optimized for throughput.
Probably not much, SVE2 has some nicer instructions, but neon already is quite solid.
> And implementing 256b SVE has different issues depending on how you do it
For in-order, and not very aggressively out-of-order cores having a larger vector length can be very useful to still get a lot of throughput out of your design. It also helps hide memory latency.
Here is a paper: https://ar5iv.labs.arxiv.org/html/2309.06865
and presentation: http://riscv.epcc.ed.ac.uk/assets/files/sc23/Short-reasons-f...
For aggressively out-of-order cores it should, for the most part, just be about decode, and some what memory latency hiding.
> 2x256b is only beneficial over 4x128b if you're limited by decode width [...] 3x256b would probably imply 3x128b which would regress existing NEON code.
I agree, that's why I don't get why people are "excited" for Zen5 to have 512b execution units, instead of 256b ones. At best there won't be a performance improvement for avx/avx2 code, at worst a regression.
Calling hwy::DispatchedTarget indicates which target is actually being used.
> 2x256b is only beneficial over 4x128b if you're limited by decode width
This is only true if we ignore more complex instructions and focus on things like adding two vectors.
> This is only true if we ignore more complex instructions and focus on things like adding two vectors.
ARM implemented a CPU that had 2x256b SVE and 4x128b NEON. Literally the only benchmarks that benefitted from SVE were because they were limited by the 5-wide decode in NEON.
Do you have an actual real-world counterexample?
However, masking can still help our VQSort [1], for example when writing the rightmost partition right to left without stomping on subsequent elements, or in a sorting network, only updating every second element.
[1] https://github.com/google/highway/tree/master/hwy/contrib/so...
I think the transition from AVX2 to AVX512 is comparable in that it provided not only larger vectors, but also a much nicer ISA. There were certainly a few projects that benefited significantly from that move. simdjson is probably the most famous example [0].
[0]: https://lemire.me/blog/2022/05/25/parsing-json-faster-with-i...
Ironically, on the RISC-V side, RVV 1.0 hardware is readily available and cheap. BananaPI BPI-F3 (spacemiT K1) is RVA22+RVV, as well as some C908-based MCUs.
simdjson specifically benefitted from Intel's hardware decision to implement a 512b permute from 2x 512b registers with a throughput of 1/cycle. That's area-expensive, which is (probably) why ARM has historically skimped on tbl performance, only changing as of the Cortex-X4.
Anyway simdjson is an argument for 256b/512b vector permute, not 128b SVE.
Having written a lot of NEON and investigated SVE... I disagree that SVE is a nicer ISA. The set of what's 2-operand destructive, what instructions have maskable forms vs. needing movprfx that's only fused on A64FX, and dealing the intrinsics issues that come from sizeless types are all unneeded headaches. Plus I prefer NEON's variable shift to SVE's variable shifts.
The sizeless headache is anyway there if you want to support RISC-V V, which we do.
One other data point in favor of SVE: its backend in Highway is only 6KLOC vs NEON's 10K, with a similar ratio of #if (indicating less fragmentation, more orthogonal).
AVX512 is all around a nice addition as JIT-based runtimes like .NET (8+) can use it for most common operations: text search, zeroing, copying, floating point conversion, more efficient forms of V256 idioms with AVX512VL (select-like patterns replaced with vpternlog).
SVE2 will follow the same route.
As to SVE though, I'd guess variable execution time makes the implementation require a bit of work. Normally, multi-cycle tasks have a fixed number. Your scheduler knows that MUL takes N cycles and plans accordingly.
SVE seems like it should require N-M cycles depending on what is passed. That must be determined and scheduled around. This would affect the OoO parts of the core all the way from ordering through to the end of the pipeline.
That's definitely bordering on new uarch territory and if that is the case, it would take 4-5 years from start to finish to implement. This would explain why all the ARMv8 guys never got around to it. ARMv9 makes it mandatory, but that was released in 2021 or so which means non-ARM implementors probably have a ways to go.
It’s not a secret that a large portion of actually technical people (not tech influencers) left Twitter for Mastodon. So the people on Twitter may be more polished turds but are turds nonetheless.
Full disclosure: I still use my Twitter for customer support because that’s all it’s good for at this point IMHO. I also don’t regularly read mastodon but it’s where the people I care about are and when I post (rarely) that’s where I do it.
That is certainly more code, but not double. You only need it for the parts of the code that are both (a) bottlenecks worth optimizing and (b) actually benefit by using the new instructions.
(Edit to remove iMovie from the list, as it has GB's of "Transitions" and "Titles" that I really should just delete)
All three examples you gave have substantial UI and other bundled assets. For example, the Docker Desktop app is about 2GB on my computer, yet included assets make up at least 1.2GB, and a further 600MB is a bundle containing the UI, which itself is about 100MB of binaries.
If you actually open those bundles (as they're called on macos) and take a look inside, you'll see that they don't even contain all of their assets, anyways, often linking to frameworks contained in ~/Library
This is a very layperson explanation, btw, but I assure you that "in modern desktop software, code is a tiny bit of the total size of the application" is a very true statement.
Yes, game assets really take it to another level but it’s been my experience that even apps without a lot of UI still have their images making up a lot of the app bundle size.
It's more comparable to x86 chips with AVX-512 and chips without AVX-512. 99% of your code is the same, but the compiler will generate SSE, AVX, and AVX-512 variants and choose the correct one based on the CPU.
If you're already supporting 2 arch it will only increase by 50 percent to support a 3rd ;-)
-----
I said "statically-linked" because the first thing that came to mind was go-lang's morbidly obese executable binary sizes, and mentally walked to iOS app sizes.
I've even seen windows binaries have multiple different versions of the same DLL inside them, and it's a well known DLL that is duplicated multiple places elsewhere.
All OSes/Apps do this but maybe a lot of Mac apps do it a little less. (I don't even have any real statistical idea how common this is with windows apps either)
- Windows DLLs don't usually have strong versioning baked into the filename. On OSX or Linux, there's usually the full version number baked in (libfoo.so.3.32.0) with symlinks stripping off version components. (libfoo.so, libfoo.so.3, libfoo.so.3.32) would all be symlinks to libfoo.so.3.32.0 and you can link against whichever major/minor/patch version you depend on. If your Windows app depends on a specific version it's going to be opening DLLs and querying them to find out what they are.
- Native OSX software (not Electron) seems to depend much less on piles of external libraries because the OSX standard library is very rich and has a solid history of not breaking APIs and ABI across OS versions. While eg CoreAudio is guaranteed to be installed on an OSX install and be either compatible or discoverably-incompatible, the version of DirectSound you're going to have access to on Windows is more of a crapshoot.
- Windows apps (except for the .Net runtime sometimes) are often designed for longevity. A couple of months ago I installed some software that was released in 1999 on my Windows 11 machine and it just worked. Bundling up those DLLs is part of why they work.
- Linux apps can rely on downstream packaging to install the necessary shared libraries on demand, generally speaking. Linux desktop apps distributed as RPMs or DEBs can "just" declare which libraries they need and get them delivered during install.
What's so newsworthy about this?
Actually probably the best thing to do is wait until the M4 machines launch then bag a good deal on a clearance M3.
When you actually need to get a new device, just get whatever the up-to-date thing is.
OK, ok, I suppose that it's reasonable to check the rumor sites to see if you should delay by a month or two. But not any longer than that.
I think Apple has been pretty good about hitting the right cadence with processor perf increases. They are making up for lost Intel time. The M6 is going to make us loose our minds. Apple is going to bring back "this is a munition" ads.
For myself, I like to think of it as applied procrastination. I could buy that new thing I want today.. but something better will come along in time, so I can afford to put it off a while longer yet..
Spot on!
Back in the nineties, Intel managed to push competing RISC architectures (UltraSparc, MIPS, DEC Alpha, PowerPC) out of the market using nothing but promises that Itanium was going to blow them all out of the water.
And apparently Apple is okay with procrastinating and cannibalizing current sales of M1, 2, 3 if it helps prevent some Snapdragon (or Ampere) sales.
sales of what
i actually can't think of a single competing product. admittedly i don't keep up with laptop news but still, i haven't heard of anything yet that can meaningfully compete with the m1 from four years ago
https://www.youtube.com/watch?v=uY-tMBk9Vx4
For me at least, the best possible outcome of this is that Windows handheld gaming devices become more power-efficient. That might be an advantage over Linux-based handhelds for a while, unless Valve decide that Proton needs to also be an architecture emulator. The chip efficiency wins must surely be tempting in this form factor.
I don't have any need for any Windows-only program.
Not sure where “procrastinating” fits in (a typo?), but as Scott McNealy once said, “If someone shows up and eats our lunch, it might so well be us.”
You may have heard of the 5-minute rule - "Will doing this take me less than 5 minutes? If the answer is yes, do it now." An adaption of that to reduce impulse purchases is - "Do I really need this product right now? If the answer is no, don't buy it."
The next one will probably have an OLED screen; so if you wait til then, your refurb M1/2/3 will be on Apple's short list of devices they don't want to support. (And you might have panel FOMO.) Or you'll have to pay the premium price for the latest model.
As a data point, I still use a 2013 Mac Pro as my primary desktop, and I've been using Sonoma on it for several months, have been able to install all Sonoma patches over-the-air on release without incident, and have only experienced a single, trivial problem: the right side of the menu bar occasionally appears shaded red, in a way that doesn't affect usability; switching applications immediately resolves the problem (the problem appears to be correlated with video playback).
Not just that, for high res stuff or modern codecs like AV1 or h265 is probably not supported at all in a 2012 device without updates for so long?
Even if support was possible it would be software encoding and even short clip it can take hours to render ?
I would happily use an older device for development a lot of dev work especially if not frontend or UI usually i can use any laptop just as a terminal, but UI or video editing I wouldn’t be able to.
I recently upgraded from a 2019 Intel Mac to a similarly-specced M3 Mac, and it really is night and day. My battery life is more than doubled - I can run IntelliJ and multiple Docker containers on battery for more than my whole work day, when before it would barely last a couple hours with that load and be slow while doing so. The fan hardly ever runs while on my Intel Mac it would run constantly.
But beyond that they are also incredibly fast and run cool. In the MacBook Air there is no fan and on the Pros they barely ever spin up in an audible way.
During work from home during Covid I was still using an Intel MBP and video conferences invariably caused the fans to kick up to the point where using noise cancelling headphones and not the built in speakers was necessary for sanity.
Battery life, portability, COOL. Like, my actual lap is no longer burning.
I think it's saving me an hour a day and the fan has never come on, the the laptop has never felt warm, and the battery life is just mind blowing.
I run docker & compilers all day. The i9 would run the fan 75% of the time and had to throttle down any time it was on battery power and it was lucky to last 3 hours on battery.
It is much, much faster, silent, and I use it for days without power. Editing 4K video is not just possible, it is a non event.
I’ll buy a little later. I’ll buy a little later!
So frustrating.
So owning a device for 6 years between age 1 and 7 will generally have a lower cost than owning a device between age 0 and 6.
For Apple products it’s generally feasible to effectively buy first hand devices aged 1+ because they’re still available for sale (at least in some retailers) after a new edition is released.
[1]: https://www.macrumors.com/2024/05/23/18-8-inch-foldable-macb...
It would require the biggest UI redesign in the history of the company to ensure every input control is at least a centimetre away from anything else.
And would require every Mac developer to absorb the cost for major updates to their apps as well.
This would almost certainly be an iPad.
And it would be possible to update macOS to enable basic touch input for unmodified apps if they wanted to.
In recent years the MBP line has been updated towards the end of the year (Oct/Nov) or early (Jan):
* https://buyersguide.macrumors.com/#MacBook_Pro_16
So if you can 'limp' along towards the autumn/winter/Christmas, then it's probably worth the wait to get the M4 (or pickup an M3 when the price presumably drops to clear inventory).
Look at real world differences between M2 and M3, it's not a massive jump at all.
I do cross platform app development and the machine is excellent for that. Glad to have it now rather than waiting months for a slightly better system
Now if that's not fun enough, their telemetry also covers mouse movements. Go ahead and watch your CPU as you spin your mouse in circles around the Teams window.
For extra fun, block their telemetry server and watch Teams bloat in RAM, to as much as your system has, as it keeps every action you take in local memory as it waits for the ability to talk to that telemetry server again.
If you're going to block their telemetry its best to fake an accept via some mitm proxy and send back a 200 code.
I do not know exactly how much this applies to iPad version, compared to their desktop apps. Mobile offers both more and less data possibilities. It's a different context.
The problem mainly occurs when I mention someone in a reply to a thread. Once I type @<name> the text input just slows down so much I can type much faster than it can render the text.
If it was a straight text box that they polled contents of _occasionally_ or after you hit 'send', it would be a much better user experience.
I'm not even saying I doubt you, I'm just curious how you ascertained this exact behavior.
From what I remember..
There are files on the disk that get updated/overwritten with pulls from the server every time it launches. Somewhere in AppData I think. A few of these are config files (with lots of interesting looking settings, including beta features).
One of the config entries specifies a telemetry endpoint (which, you _could_ figure out with a network tracing tool but there are a ton of MS telemetry endpoints your machine is probably talking to. Best to just grab the one explicitly being used from the config like this). I forget the full name of the setting but the name pretty clearly indicates its for telemetry, and the file is clearly a config file. If you can't find it just by browsing the structure, try a multi-file search tool and look for 'telemetry' or URL/hostnames.
You can't really change the value on disk and make it just take effect from there, since it gets downloaded from the server and overwritten before Teams loads. There might be some tricks you can do locally to persist the change but nothing seemed to work for me. You could override response from server via mitmproxy but that requires finding where it comes across the wire at launch time and then building a script/config to replace it.
Anyway, you can block that telemetry endpoint from a firewall and see your memory bloat. Or you can intercept that endpoint in any mitm proxy. I went with this [mitmproxy](https://mitmproxy.org/). From there you can capture the content it sends to the endpoint, or even change the response the server sends (Teams just seems to expect a 200 code back).
The telemetry data itself is some kind of streaming event format. I think I even found documentation on the structure on some microsoft website, so its likely a reused format.
It's pretty straightforward.
I couldn't spend too much time on it and now it's not something I even use, but some cool things you might want to try if you dive deeper into this:
- Overwrite the config file as it returns from the server, to turn on EU data protection, change various functionality you're not supposed to, or flip some feature flags.
- Figure out if there's a feature flag or even other overwrite to fully disable the metrics so they aren't even collected, from anywhere in the app.
- Intercept telemetry, return an 'OK' response and drop the data from telemetry, or maybe document what they collect more definitively if you think there's interest somewhere. This keeps your privacy but doesn't really do anything for performance.
- Interfere with the data before actually returning it, maybe try playing with event contents and channel/user indicators. Microsoft probably won't like this if they notice, but it's unlikely they'll even notice.
1. https://portswigger.net/burp/documentation/desktop/getting-s...
Edit: and that reminds me I should probably run this test on new Teams, where it now uses the built in WebView2
Please tell me this is a satire piece...
Never underestimate the ability of shitware vendors to make supercomputers feel slower than an 8086. These days it usually involves JavaScript, HTML, and CSS.
A glorified IRC client should run in under a megabyte of memory.
I believe it is actually hitting the server to update the online/away status light for every single message in a conversation. If you turn off all the status update stuff in the settings then the software speeds up dramatically. Another thing you can do is find the folder where it caches everything and just trash the entire thing. Somehow, they’ve managed to make caching slow everything down rather than provide a speed up.
It’s exactly the same but then you can have to open twice and drain your resources much quicker.
Teams installed directly from the App Store or Microsoft on any device I own (high end windows/mac/ipad) devices are all terribly slow
I haven't done async JS, but I've done enough async elsewhere to know that language cannot work around bad design.
Work started off using Google talk. I am expecting something to replace google chat like they did talk and hangouts.
Maybe someone should normalise giving developers crappy laptops to develop on.
(Has anyone done a deep dive into Teams to explain what on earth is going on? I mean, if VSCode can be fast despite its underlying architecture, surely something could be done about Teams?)
Then the developers will complain the hardware is unusable to do their job even though that this was a supercomputer back in the day. Then you say "No, it's the software please fix it."
Some days it makes me extra motivated to make the code I write fast and efficient; other days I want to give up entirely.
Would you give a professional smith a low quality hammer? Would you give a chef a blunt knife?
Why do corporations give out garbage hardware to their smiths and chefs (+ enforce their use via policy) ?
One as to hope that the same perfomance lens will now be turned on the mobile apps.
Oryon will still probably beat x86 designs massively in performance per watt which is pretty much the most important metric for most people anyway (as most people use laptops).
EDIT: your username `dragonelite` is quite interesting. You joined 2019, but the coincidence is fascinating.
Lmao, and people say Geekbench isn't biased towards ARM
It was a super great idea because it allowed recompilation on the App Store to take advantage of new instructions.
Updating apps to use new vector instructions is far more complicated than upgrading to a new compiler version and having it magically get faster.