Now they have the memory chips "inside" the "CPU", aka: system in a package.
Big Tim is way ahead of the game when it comes to ruthless product targeting and segmentation.
https://www.servethehome.com/intel-xeon-max-9480-deep-dive-i...
Alas the 1+4 Skylake+atom were all pretty meh chips & Intel hasn't tried again. Conceptually I really.liked it, and it was doing great against the Qualcomm 8cx competitor for the tablets/2-in-1s it was in, but people wanted more.
I really hope the hbm memory on-die trickles down into pro-sumer processor chips eventually too.
Increasing memory speeds have had only a marginal impact on common workloads (not AI) in the PC World - this has been a well discussed topic. Why would it be different on the Mac?
Memory Latency, on the other hand, can have a huge impact on performance. That's the reason chips have had steadily increasing cache sizes over the decades along with the introduction of additional caching layers.
What Apple's M series improves by putting the memory on chip is both bandwidth and latency. The latency, however, is what will impact app performance more than anything else.
Latency improvements are the reason memory controllers were moved from motherboard northbridges into the CPU IC. Shortening that distance means a lot to shortening latency.
It's not an unreasonable hypothesis that this is because of on-package memory which means you can run with tighter thresholds since your memory traces are shorter. Their unified memory architecture also means that GPU is way more effective since you never need to copy memory from the CPU to GPU. That's probably not a huge deal for normal apps, but having rendering the screen not competing with any app resources for memory probably isn't nothing.
Neither the CPU nor the GPU can fully utilize the available memory bandwidth [4] The insane memory bandwidth is to also support the other coprocessors like the media engine & to ensure that you can do a bunch of tasks in parallel without them contending with each for resources. Those workloads do benefit greatly from memory bandwidth & the "artist" community is hugely important to Apple.
The reason M is overall faster is more because of things like better ILP (e.g. fixed-width RISC encoding vs compression-like CISC means easier to keep the pipeline fed). Apple software for the most part is better engineered for performance & the M chip is optimized for that performance profile (e.g. C# and Java vs Swift/ObjC - Rc is pretty conclusively faster & requires less memory overall not to mention that the M chip has specific optimizations to further reduce the traditional cost of Rc that they have to pay because they don't typically distinguish Rc from Arc as Rust does). The final truth is that Apple throws a lot of money at this problem as well & buy up the latest processing node to stay ahead of the competition.
[1] https://www.linkedin.com/pulse/apple-m-cpus-probably-much-fa...
[2] https://www.anandtech.com/show/16252/mac-mini-apple-m1-teste...
[3] https://www.anandtech.com/show/17047/the-intel-12th-gen-core...
[4] https://www.anandtech.com/show/17024/apple-m1-max-performanc...
These have little to do with memory packaging and a lot to do with cache and prefetching architecture.
It does have one of the lowest latencies for atomics[1][2] which is something Swift cares about, but the overall core design just as much benefits other, less indirection and synchronization heavy languages (side note: Swift and Java are closer with each other in defaulting to virtual dispatch, where-as C# sits in the middle of the road with static dispatch by default but some code heavily using interfaces and virtual methods which do make such calls virtual to an extent (JIT has DPGO to devirtualize, AOT can uncoinditionally devirtualize in certain scenarios too, similar to what Swift's WMO does)).
[0] https://gofetch.fail/files/gofetch.pdf
[1] https://dougallj.github.io/applecpu/measurements/firestorm/D...
[2] https://dougallj.github.io/applecpu/measurements/firestorm/C...
Not sure how it works on the iMac/MacPro
In case of intel macs, the soldered RAM should have made the device and RAM options even cheaper, as the BOM is simpler. The only real argument against soldered RAM was DDR5 in SODIMM being bad - which LPCAMM2 fixes.
If they didn't use it specifically to price gouge, the soldered RAM wouldn't have been a problem. 64GB of RAM costs practically nothing at market prices, so they could honestly have had a single SKU with max RAM without notable price change over entry level - heck, maybe the BOM reduction would even sponsor it.
... But then they'd lose their mechanism to drain companies with deep pockets for necessary upgrades. Being sensible and fair is not profitable.
https://i.imgur.com/Y3PLp33.jpeg
Anyway results are what matters, and the module in the OP has bandwidth figures half way between the M3 and M3 Pro while retaining upgradability.