Apple Silicon Mac Pro uses PCIe switches to support all slots
social.treehouse.systems
social.treehouse.systems
Finally Woz’s design for the Apple II color graphics was based on unified memory.
Dramatic shifts in technology sometimes obviates all our past learning and experience.
PCIe is a tree-shaped network (root is the root hub in the CPU, inner nodes are switches, leafs are devices) with point-to-point links (called lanes) that transmit packets. Bandwidth is determined by the number of lanes that a device's path up to the root hub has at the narrowest node.
Switches are far more complex. The protocol specifies addressing, bandwidth reservations, flow control, priorities and message ordering requirements, which the switches have to handle.
Not a switch thing per se, but also relevant for "single transistors won't do": On the physical layer, each port/link is also relatively complicated to initialize and operate. The transmission encoding, due to the relatively high speeds involved, requires a learning phase where both ends measure the link characteristics and adjust their (relatively complex) signal processing accordingly. The physical layer also handles different link speeds for down-/upward compatibility.
This may explains how they achieve the graphics blitting in the Vision Pro.
I would think this would lead to bottlenecks for instance if you were trying to write to two different 16x PCIe SSD RAID cards simultaneously.
As an extreme example, when and where it was/is still viable, cryptocurrency mining operations use external PCIe cases with lots of slots, attached to only a few or even just one lane per case. For their purpose this was still viable, since mining uses very little I/O from/to the GPU, essentially only a small block of setup data into the GPU ("try hashing this block, vary those bytes within that range, report when result matches this pattern", a few kB) and a few bytes of results out ("found a match, varied bytes were as follows").
You are sending the exact same data to both GPUs. If the the PCIe switch and GPU drivers support multicast, then you would actually get the full 16x bandwidth to both.
Which is actually a huge advantage over the most common multi-gpu setup, which is forced to statically split the 16 PCIe lanes into two independent 8x slots. Even if you were playing a game that only used one GPU, simply having the second gpu plugged in meant your total bandwidth to each gpu was limited to 8x speed.
I really wish you could find x86 motherboards with PCIe switch chips... but switch chips are really expensive so nobody makes them.
Not sure how this affects the use case of multiple video capture cards writing to fast internal storage that I see mentioned here and there. Anyone close to that industry around to comment?
Edit: i think I’ve seen mentioned that the PSU in this thing is “only” 1400W. That only covers one modern nvidia gpu right? :)
The graphics and processing speed of the Vision Pro makes me wonder if the GPUs are truly needed.
Most of the cards for the M2 Mac Pro are used to access the analog world. It seems that there is power inside the die and if you need an analog signal use a PCI-e port.
For GPU compute, I'm sure it can't beat 15 kW of nvidia dedicated cards.
But then you don't want 15 kW of cooling on your desk.
I'm mostly curious as to what this work load consists of?
but when I read the number 8k encode and decode streams 15 kW represents
How much of the power is just getting the data in and out of the card to the processor. Apple Silicon doing the same work on die is going to use less power. I don't doubt that there are faster processing units.
- Bad: bottlenecked access to the CPU and its memory
- Good: Bypass the CPU on device to device transfer, in the style of NVidia's GPUDirect (https://developer.nvidia.com/blog/gpudirect-storage/)
Obviously, but that does not explain why CPU hot swap is inherently expensive. Engineering effort is required to make any CPU.
Why must the engineering effort specifically relating to CPU hot swap inherently cause architectures that supports CPU hot swap significantly expensive compared to architectures that does not?
As a lay man without any experience with actual semi conductors outside of university, I would have thought that this would mostly be something that needs to be designed and specified once and then it could be mass produced. Clearly it will take some effort and you would have to spend some gates on it in the chip design.
Considering all the other stuff that are in modern CPUs and chipsets, I would have thought it would be possible to do rather cheaply if mass produced.
———
Because if you don't, sure, you didn't technically power down your machine, you just fucked up your entire stack, stack pointer goes to the wrong place, checks based on hardware are now wrong... In effect, you powered down the machine because your OS will just stop working, your running processes are fucked, etc.
A good programmer familiar with the target OS and inner workings of PCIe could deliver a working prototype within a week, production ready within less than a year.
> A good programmer familiar with the target OS and inner workings of PCIe could deliver a working prototype within a week, production ready within less than a year.
Could you please explain why it'd be normal to take a year to go from prototype to production in this case? What would happen in that one year, starting from a working prototype?
and yet, Apple charges for it as if it was a real workstation equipment rather than a phone in a massive enclosure.
Not really sure how many ppl would use this but if apple switches to Type C connectors they’re more likely to strut it at the high end than just just mention it in passing.
The use case isn’t that clear to me. Perhaps a photographer, for a phone with a couple of TB of SSD.
IMO, eventually, phones will replace low end desktop/laptop computers.
Someone (probably Apple, since their CPUs are head and shoulders above the competition) will have to do the work of designing a mode switching UI, but once that's done, they'll really be able to cut into desktop volume.
I'm sure the early devices will be... awkward to use - mostly limited to the browser[1].
But after a few iterations, I'm sure it'll mature into something that can stand on its own.
No idea when it'll happen, but I do think that economics will push it into being sooner or later.
---
1. Presumably people would also be able to use other phone apps, but a lot of those look weird when scaled up. And transiting from touch to mouse/trackpad is not straightforward.
The trivial way I disagree is that I believe people don’t want this: they want to be able to use their phone while “computing”. Perhaps AirPods are enough for that (use Siri to make a call while otherwise mousing) and thus I’m wrong.
The big way I disagree is that the phone is already many people’s only internet terminal and and for most has taken over almost all uses of a “computer”. Even I use my bank’s phone app when sitting in front of my computer because it has features the web interface doesn’t.
It's also not really new, and has been tried many times in the past. I do think we finally have all the requisite technology in place though. BT peripherals mostly work, phones CPUs are powerful, phone apps are plentiful, etc...
The trouble with the future is that, even if you know that it's inevitable, you may be too early for that to matter.
There are a huge number of dot com busts that later became successes (different company, though) because we have the needed technology today.
I still think that is true and will eventually be building my own on a Mac Pro. I will be combine various sources into a single context that will represent the building. VisionOS has changed a lot of my thinking. I would like to be able to access the hand-tracking separate from goggles.
I mean you can't style yourself a workstation if the lights don't dim when turned on
Apple silicon is made using mobile libraries, same as their iPhone and iPad chips. Even though they're made on these dense libraries, the chip is still very big, since they need to include CPUs, GPUs, NPUs, and all of their fixed function hardware.
All of this puts a limit in terms of frequency and core count. You will find that other server hardware in this price range significantly outperforms Apple silicon , save for those video workloads they're missing fixed function hardware for.
M2 Ultra has some 134B transistors, when you compare with an equivalent chip like MI300 (146B), you find that really, in terms of raw compute, there is no comparison. We will have more details on MI300 today, watch out for the news cycle.
This is an excellent read in terms of what difference using dense vs performance libraries can make - https://www.semianalysis.com/p/zen-4c-amds-response-to-hyper...
I think the power envelope, not the price tag, is the better metric when comparing how server-y or phone-y the silicon it.
> "M2 Ultra has some 134B transistors, when you compare with an equivalent chip like MI300 (146B), you find that really, in terms of raw compute, there is no comparison. We will have more details on MI300 today, watch out for the news cycle."
And unsurprisingly it will certainly come with an incomparable TDP.
You do know that the Mac Pro is targeted specifically at those workloads.
If you don't have those constraints it is easy to just "copy and paste" more cores and naturally this will have a higher power envelope, but in essence this is a phone tech that you pay extreme premium for.
M2 Ultra has 800GB/s, which sounds like a lot until you remember a Xbox Series X has 560GB/s and costs $499. Or spend $800 and get a 7900 XT which has 800GB/s.
Think of the Ultra more of a GPU that happens to have CPU built in.
That's not really the pejorative that you think it is.
Intel crushed the Unix workstations with volume - it was more economically efficient to make a workstation version of a desktop chip than to design one from scratch that wasn't supported by the desktop's volume.
Well, Apple is doing to Intel what Intel did to HP, Sun, IBM, DEC, and SGI: using their phone (and watch, etc.) volume to compete with Intel.
Perhaps Intel (and AMD) will stave them off. But if history is any guide, they'll be paddling upstream to do so.
Ignoring the no true scotsman argument, I'm not sure this is even an insult. Apple improved their phone CPUs to the point where they compete very well in the desktop space and in some workstation (for any meaning of this term) environments. It's what Intel did to traditional workstation vendors long ago, and is classic Innovators Dilemma. Typically that happens from a small startup and not some industry leading company though. Shame on Intel and AMD for not seeing this coming, since it had been telegraphed for years.
Why worry it someone calls it a workstation. That has to be one of the world's most boring technical arguments.
The unified memory is great for LLM style tasks of course and it's one of the few ways to get a GPU setup with over 100GB for less than $100k.
PS Not sure how the comparison with a 20 year old system applies though ;)
One instance is that PCI lanes didn't exist when the term was conjured up therefore can’t be part of inherent definition.
The mac pro has 64 lanes acroos 6 slots. PCI-e external expansion is another 8 ports of expansion limited to 40 GB/s. Unfortunately the concept of lanes doesn't seem to exist for external PCI-e ports so its hard to judge if those would count as more lanes than a normal desktop.
The 20 year old workstation is to demonstrate that the slab form factor has been called a workstation from the beginning.
It's has always been a subjective term and I believe it has more to do with output than consumption of resources during a compute cycle. There is no reason to try to dispute some claim to the term.
https://www.sonnettech.com/product/legacyproducts/rackmacpro... https://www.myelectronics.nl/us/mac-studio-empower-station-r...
At least the new case is available with rails, but cooling and management are still crap.
But the 2019 mac pro did have an official rack mount variant of the case.
Those are certainly users with I/O needs, but very different ones from the needs of a server.
On the other hand the comparison between Apple Silicon CPUs and "server-grade CPUs" has always been much less favorable for Apple. The reason is that the higher IPC of Apple CPUs does not matter for server applications. The server CPUs use the same low clock frequency as the Apple CPUs, resulting in the same low power consumption. An Intel or AMD core is slower than an Apple core, but that is compensated by more cores, and more cores with lower IPC have a better performance per watt than fewer cores with higher IPC.
Therefore, until recently the Apple Silicon CPUs had a single advantage vs. "server-grade CPUs": they were made in various "5 nm" process variants, while the best "server-grade CPUs" were made in various "7 nm" process variants.
Since the end of last year, and especially after the launch of today of AMD Bergamo, the existing Apple Silicon CPUs do not have any advantage for server applications a.k.a. multithreaded applications.
Only after the launch of the next Apple Silicon CPUs, made in a "3 nm" process, they would become competitive again. However that does not matter anyway, because Apple has abandoned a long time ago the idea of competing in server applications and for now they are not interested in such applications.
Whenever the intention is to praise the good performance per watt of Apple Silicon CPUs, they should always be compared with laptop or desktop CPUs, never with "server-grade CPUs", because in the latter case the numbers are very different.
I’m not sure how to interpret this. Of course server applications care about IPC. Maybe moreso! The only difference I can imagine is that server applications are more likely to be memory-bandwidth bound, which does indeed make IPC irrelevant, but also plays in favor of the high bandwidth architecture of the M- series.
A high IPC is good for interactive applications that cannot be parallelized, so the latency between request and reply is determined by the product between IPC and clock frequency. Intel and AMD can achieve a lower latency than Apple only by raising the clock frequency, which is paid by an increase in power consumption that exceeds a lot the increase in performance.
So choosing to design a CPU with a higher IPC has contradictory effects on its efficiency for throughput-oriented (server) and for latency-oriented (personal computer) applications.
But run-of-the-mill server workloads are more like database applications (get query, load data from mass storage, do some processing on it, send results to network), HTTP and similar servers (get request, load data from mass storage, do some processing on it, send results to network), file servers (load/store data from network, only a little processing), terminal servers (wait for user interaction on network, process/load/store a bit, shovel out screen updates). Lots and lots of I/O, not so much computation that would care about IPC.
Intel top end is Xeon. i9 is gamer CPU, more or less.
It is really starting to look like either COVID and new Campus architecture had seriously crippled engineering capabilities of Apple, or if not all the capabilities are redirected into futuristic projects such as cars and Vision VR goggles. They could have made Pro a cluster supercomputer on a backplane just as a flexing and a nod to NeXT, but didn't. They just spread out Studio motherboard into existing Pro case. Sony would do more.
Apple is crippled because they didn’t prioritize a niche use for the sole purpose of “flexing”?