Apple M1 Ultra
apple.com
apple.com
If Pytorch becomes stable and easy to use on Apple Silicon [0][1], it could be an appealing choice.
[0]: https://github.com/pytorch/pytorch/issues/47702#issuecomment... [1]: https://nod.ai/pytorch-m1-max-gpu/
5-10 years ago they were still serious about open standards, like OpenCL.. Now it's all locked in.
1. https://techguided.com/best-rtx-3090-gaming-pc/#:~:text=With....
That 3500 is for a DIY build. So, sure, you can always save on labor and hassle, but prebuilt 3090 rigs commonly cost over 4k. And if you don't want to buy from Amazon because of their notorious history of mixing components from different suppliers and reselling used returns, oof, good luck even getting one.
(I've been accused of overuse of acronyms, but that one's rare!)
EDIT: I’m quite confident this is not at all an exaggeration. Unless you have put together PCs for a living. $100/h (total employment cost, not just salary), 1-2 hours of actual build & setup, 8 more hours of speccing out parts, buying, taking delivery, installing stuff and messing around with windows/Linux (I’ve probably spent 40 hours+ in the past couple years just fixing stuff in my windows gaming pc. At least 1 of those looking for a cabled keyboard so I could boot it up the first time, ended up having a friend drive over with his :D)
You have to spec out the parts, ensuring compatibility. Manage multiple orders and deliveries. Assemble it. Install drivers/configuration specific packages.
All of these things are easier today than ten or twenty years ago - but assigning it to a random mid-level engineer and I'd set my project management gamble on half a day for the busiest, most focused engineers least likely to take the time to fuss over specs, or one day for the majority.
ofc. to get to $1000 for that they'd still have to be on $230k to $460k.
In my case I have built a couple PCs before, but it was so long ago that I'd have to re-learn which retailers are trustworthy, what the new connection standards are these days, etc. It's just not worth it to me to spend a dozen hours learning, specing, ordering, assembling, installing, configuring, etc to save a few hundred bucks.
A senior engineer in the Bay can easily pull down $400k/year in total comp, which is $200/hour. The rule of thumb I've always heard is that a fully-loaded engineer costs roughly 2x their comp in taxes/insurance/facilities/etc.
When someone costs the company north of $3k/day, it's cheaper all round to just plonk a brand new $6k MacBook Pro on their desk if they have a hardware issue.
So eventually you still have to "replace everything" to upgrade a PC.
You could also put a 4790k, 16gb of ddr3 and a modern gpu in that system to get a perfectly functional gaming system that will do most titles on 1080p high. Though admittedly we've passed the point where that's financially sensible vs upgrading to a 12400 or something as both devil's canyon CPUs and ddr3 are climbing back up in price as supplies diminish
DDR4 was released in 2014, which would suggest you purchased your mobo two full years after DDR3 was already deemed legacy technology and being phased out.
Also LGA1150 was succeeded by LGA1151 in 2015, which means you bought your mobo one full year after it was already legacy hardware.
Point being, PC-building makes it easier to replace and repair individual components, but in time, upgrading to newer generations means spending over 50% of the original cost on motherboard, CPU, PSU, RAM. Not too different than dropping $3K on a new Mac.
[1] https://web.archive.org/web/20101219085440/http://www.xbitla...
It means the hardware was purchased after it started to be discontinued.
It's hardly a reasonable take, and makes little sense, to complain how you can't upgrade hardware that was already being discontinued before you bought it.
> DDR3 and LGA1150 were not deemed "legacy" the day DDR4 and LGA1151 motherboards entered the market.
I googled for LGA1150 before I posted the message, and one of the first search results is a post on Linux tech tips dating way back to 2015 on whether LGA1150 was already dead.
And you purchased the Mobo one year after that.
Also, we must have a different interpretation of "discontinued", because DDR3 and LGA1150 were still produced, sold, and dominated sales for way long after I bought that system. At the time (and for the next 1-2 years), consumer DDR4 was a luxury component that most no existing hardware supported.
To do CPU upgrades you eventually have to replace the motherboard but you can keep using whatever your GPU/storage/other parts is. Sometimes that also means a RAM upgrade but it's still better than the literal nothing of modern Macs.
Performance per watt? I could see that being disrupted, but an iGPU in 2022 will be orders of magnitude less powerful than a dGPU, if wattage is ignored.
If Nvidia were on the same node and increased die space to match M1 (ignoring the CPU portion of the die size), they would then be able to run at a lower clock with more compute units and probably match the TDP discrepancy.
An iGPU isn't necessarily slower if the system ram is fast, and M1 was one of the first consumer CPUs to move to DDR5. 3090 has 936.2 GB/s with GDDR6X, M1 Ultra with DDR5 memory controllers on both dies gets 800GB/s.
For the record: I was the first M1 recipient (temporary 16gb MB, stock issues). I needed an Intel MBP because Rosetta ain't all that. I opted for, and was upgraded to, the 32gb M1 MBP. I chose M1 over Intel because it was unbelievably faster for the form-factor. My original comment does not concern laptops. My PC is orders of magnitude more powerful.
TDP is physics. You all might perceive Apple as perfect and infallible and all so lovely, but physics is physics.
I use AMD, not NVIDIA. And "what if" is irrelevant. It's like intentionally neutering Zen2 by comparing it to Intel single-core (as was done all the time). The reality is absolute, not relative. Comparing effective performance, not per-TDP, is what matters to the user. And my network/gpu/audio drops on both my 16gb M1 MB and 32gb M1 MBP under load.
Seriously not buying that "Apple can do nothing wrong" bias.
Take those #Ff6600-colored glasses off. The M1 has unbeatable value proposition in a pretty wide market, but Apple couldn't be further from a universally good machine.
Here in Australia, 3090’s go for close to 3k on their own.
20-Core CPU 48-Core GPU 32-Core Neural Engine
64GB unified memory
1TB SSD storage¹
Front: Two Thunderbolt 4 ports, one SDXC card slot
Back: Four Thunderbolt 4 ports, two USB-A ports, one HDMI port, one 10Gb Ethernet port, one 3.5-mm headphone jack
A$6,099.00But even ignoring the FE series, the prices have already crashed massively, you can get a 3080 AIB for less than £1000, and 3090s frequently appear around £1500-1600.
But you’re comparing apples to oranges, because the real advantage of M1 chips is the unified memory - almost no CPU-GPU communication overhead, and that the GPU can use ginormous amounts of memory.
Since you mention ML specifically, looking at some benchmarks out there (like https://tlkh.dev/benchmarking-the-apple-m1-max#heading-gpu & https://wandb.ai/tcapelle/apple_m1_pro/reports/Deep-Learning... ), even if the M1 Ultra is 2x the performance of the M1 Max (so perfect scaling), it would still be far behind the 3090. Like completely different ballpark behind. But of course there is that price & power gap, but the primary strength of the M1 GPUs seems to really be from the essentially very large VRAM amount. So if your working set doesn't fit in an RTX GPU of your desired budget, then the M1 is a good option. If, however, you're not VRAM limited, then Nvidia still offers far more performance.
Well, assuming you can actually buy any of these, anyway. The M1 Ultra might win "by default" by simply being purchasable at all unlike pretty much every other GPU :/
edit: Of course the M1 max only shipped in laptops, so... who knows.
I’d assume that’s what most of the chonk is about, no?
> Based on the thermal and power consumption characteristics of previous chips I would not be surprised if say ~120W is the max power draw of this thing.
The Max could be brought up to 90W or so.
https://www.theverge.com/22967776/apple-magic-mouse-charging...
If you think this is still a problem, you haven't used any recent Macs. The current MB Air and MB Pro both run very cool even under prolonged heavy loads.
Apple's management of any and all heat issues has been far better than any competitors for a while now.
Only if you define "for a while" as "since a year ago with the introduction of the M1".
Apple refused to make a thicker laptop or one with better ventilation to adequately cool the CPUs & GPUs they were sticking in them. They were among if not the worst of them all at handling the heat of the components they were using. Until the M1 Pro & Max rolled around, anyway, and suddenly they got thicker, with feet that raise it farther off the desk, and absolutely massive amount of vents all over 3 sides of the machine. Curious timing on that...
And of course, Apple has made huge progress since then (M1, better thermal designs, new fan designs which are quieter and more efficient) whereas PC makers have made basically zero progress.
The "huge progress since then" was a year ago.
Your timeline is a little bit made up.
Back when that M1 MAX vs 3090 blog post was released, I ran those same tests on the M1 Pro (16GB), Google Colab Pro, and free GPUs (RTX4000, RTX5000) on the Paperspace Pro plan.
To make a long story short, I don't think buying any M1 chip make senses if your primary purpose is Deep Learning. If you are just learning or playing around with DL, Colab Pro and the M1 Max provide similar performance. But Colab Pro is ~$10/month, and upgrading any laptop to M1-Max is at least $600.
The "free" RTX5000 on Paperspace Pro (~$8 month) is much faster (especially with fp16 and XLA) than M1 Max and Colab Pro, albeit the RTX5000 isn't always available. The free RTX4000 is also a faster than M1 Max, albeit you need to use smaller batch sizes due to 8GB of VRAM.
If you assume that M1-Ultra doubles the performance of M1-Max in similar fashion to how the M1-Max seems to double the gpu performance of the M1-Pro, it still doesn't make sense from a cost perspective. If you are a serious DL practitioner, putting that money towards cloud resources or a 3090 makes a lot more sense than buying the M1-Ultra.
Apple Silicon (including base M1) actually has great FP16 support at the hardware level, including conversions. So it is wrong to say it only supports FP32.
> The Core ML runtime dynamically partitions the network graph into sections for the Apple Neural Engine (ANE), GPU, and CPU, and each unit executes its section of the network using its native type to maximize its performance and the model’s overall performance. The GPU and ANE use float 16 precision, and the CPU uses float 32.
Also, this exploration (https://tlkh.dev/benchmarking-the-apple-m1-max#heading-neura...) reports the 5.1-5.3 TFLOPS FP16 ballpark performance.
Thank you for recalibrating me to actual reality and not Apple Reality (tm)
https://www.techpowerup.com/64683/nvidia-admits-to-selling-f...
Apple specifically: https://support.apple.com/en-us/HT203254
Anyway, it was still my longest lived Laptop. My Sony VAIOs were great but I liked that Mac better.
Even though it's not an apples to apples comparison, keep in mind that a 1x32GB DIMM sells for less than 150$, and you can buy 1TB SSDs for less than 100$.
> keep in mind that a 1x32GB DIMM sells for less than 150$
Keep in mind that M1 Pro/Max/Ultra is LPDDR5 6400 (https://www.anandtech.com/show/17024/apple-m1-max-performanc...) connected by a 512 bit memory controller.
Whereas is a kit of 2x 32 GB LPDDR5 4800 (I could not easily locate a quote for 1x 64Gb LPDDR5 4800 DIMM, leave alone 6400) retails for USD 548 (https://www.newegg.com/crucial-64gb-288-pin-ddr5-sdram/p/N82...).
I could not locate a reliable source on the type of the SSD employed in M1 Pro/Max/Ultra, so I will refrain from remarking on the comparison.
And what WD advertises for the Black SN850: https://www.westerndigital.com/products/internal-drives/wd-b...
And what Seagate advertises for the FireCuda 530: https://www.seagate.com/products/gaming-drives/pc-gaming/fir...
And what Gigabyte advertises for the Gen4 7000s: https://www.gigabyte.com/Solid-State-Drive/AORUS-Gen4-7000s-...
etc...
They aren't $100 for 1TB, no, but a lot of them are around $150. Which would be a lot less than +$100 to go from 512GB to 1TB, too. It's $40 to go from the 500GB SN850 to the 1TB SN850, for example.
When configured to ensure data integrity in the case of power loss (more important in this new M1 Studio machine unless it comes with integrated battery), then it's a lot worse.
They're overpriced for sure, and that's the only reason the M1U pricing looks equivalent rather than exorbitant
* Much higher than MSRP, which was reduced compared to the previous generation because that attempt at raising prices killed market demand for said previous generation.
* No longer affordable by the traditional customer base but sustained by a new market with questionable longevity in its demand
Take your pick.
Sure, in a pure rational economic sense the market price has risen because supply has fallen at the same time a new market of buyers became very interested in the product, but we're talking consumer expectations and historical trends here, not the current price in a vacuum.
I wasn’t a DVer, but they’ve always been available if you were willing to pay a scalper. The only thing that has changed is that more retailers find it appropriate to rip off their customers. It’s kinda like a liquor store I’ve done business with for years now wanting $3,800 for a bottle of 23yr PVW. MSRP is $299.99. Should the owner not be able to make extra profit on it, no, of course they should. But >12x MSRP is just predatory imho.
edit: just found out - it's NBB:
https://www.notebooksbilliger.de/
There's probably a German discord somewhere to have alert for drops.
And that is expected, a lot has to be reserved for USB devices.
No, not really. You can try adding sound dampening, like BeQuiet does, but it's not as effective as just having more lower RPM fans (although it does help with coil whine). But Apple has historically never used sound dampening, and this doesn't look like it changes that. With how "open" it is anyway (the entire back being just a bunch of holes), sound dampening wouldn't be all that effective.
I'm not really sure what you think the heavier heat sink has to do with either the fan RPM or the noise profile. The bigger heatsink is if anything evidence of larger, lower-RPM fans. They're using more surface area, so they can spread the air movement out over larger fan blades. Which means they don't need to use as high an RPM fan.
There is raw performance, but there is also performance per watt, availability and scalability (which is both good and bad - M1 is available, but there is no M1 Ultra cloud available). If you want a multi-use setup, an RTX makes more sense than most other options, if you can get one and at a reasonable price. If you want a Mac, the M1U is going to give you the best GPU. In pretty much all other setups there are so many other variables it's hard to recommend anything.
Even assuming literal 24hr/day usage at a higher, "factory overclocked" 450w sustained, at a fairly high $0.30/kWh that's $1200/yr. Less than half the retail price of a 3090. And you can easily drop the power limit slider on a 3090 to take it down to 300-350w, likely without significantly impacting your workload performance. Not to mention in most countries the power cost is much less than $0.30/kWh.
At a more "realistic" 8 hours a day with local power pricing I'd have to run my 3090 for nearly 10 years just to reach the $2000 upgrade price an M1 Ultra costs over the base model M1 Max.
At $0.36/kWh - card alone @450W ~running cost ~= RRP 3090 $1,499
Yes, can power it down to be more efficient, however, that effectively agrees with the previous comment that PPW matters.
If you are doing it solo, with "just the hardware you happen to have", it matters a bit less. If you are doing it constantly to make money, and you need a lot of power, buying a one-person desk-machine makes no sense.
Watts are dollars that you'll continue spending over the system life. It matters because you can only draw so many amps per rack and there will be a point when, in order to get more capacity, you'll need to build another datacenter.
These machines are commonly used by professionals in industries like movie and music production. They don't care what the power bill is, it's insignificant compared to the cost of the employee using the hardware.
Oh... They do. At least, they should. If a similar PC costs $500 less but you spend $700 more in electricity per year because of it, at the end of the year, your profits will show the difference.
I ran a visual effects company a decade ago. We bought the fastest machines we could because saving time for production was important. The power draw was never a factor; a few catered lunches alone would dwarf the power bill.
Honestly asking because I’m kind of out of the nvidia loop at the moment.
Legally: I assume no cloud provider will assume the legal risk of telling their customers "and here you have to break the EULA of the NVIDIA driver in that way to use the service". In Europe where the legal environment is more focused on interoperability, this might not be as much of a problem, but still it may be too much risk.
The silly thing about it is that most of the special engines can now be flashed into an FPGA which is becoming more common in the big clouds so special offload engines aren't that big of a deal when they are missing. So in some cases you can have your cake and eat it too; massive parallel processing and specialised processing in the same server box without resorting to special tricks (as long as it's not suddenly getting blocked in future software updates).
Don't know if it's the same, these days, but when I was designing electronic stuff, we were always told to spec the power supply at twice the maximum draw.
Once again we get a Mac Semi-Pro Mini (seems like the Studio is more like a replacement for the Trashcan) that their marketing implies is maybe as good as a Mac Pro but is obviously not. It does look a lot better this time around - at least it has more ports :-D
Interested to know what you think a reasonable PSU would be for A machine that was consuming close to 200W for processing...
Won't even get into that it can't run most things you'd want that kind of hardware for.
No one plays games on a Mac.
And it has nothing to do with GPU performance but rather the fact that the audience simply isn't interested in gaming on it and so there is no economic incentive to target them.
So the GPU performance that matters to Mac users and is relevant to Apple is not games but rather content creation, production etc.
Video content creation specifically is where it mostly achieves what the graphs indicate, and that's mostly a "well yeah, video decoder ASICs are really efficient. I'll take things we already knew for $100, Alex"
If the software is optimized, the graphics hold up fine.
My wife does (I play games on Linux).
I have friends who own Macs who reluctantly dual boot to Windows just to play some games -- they would completely ditch Windows if they could just play every game on Mac.
I see there are Mac games on Steam.
All of this points to the situation being more nuanced than "no one plays games on Mac".
According to the steam hardware survey, Windows is 95%, Macos is 4%, and Linux is 1%. And to dig deeper you'd need to see what games that 5% of the non-Windows is playing - are they simpler games that don't need graphics acceleration (e.g. puzzle games, roguelikes, etc) or ones that do?
My desktop is an Intel NUC running Ubuntu, and yeah I play games on it. Slay the Spire, Spacechem, even some older MMOs like DDO or LoTRO (which run but at 15-20 fps since that system just has Intel Iris). I'm unable to even start many others (e.g. Grim Dawn) due to not having dedicated graphics.
So yeah it's nuanced but lots of games that need a graphics card don't run or even display on that system.
That's why I have a windows gaming system too. I'm realistic, the market just isn't there. I used to have a Mac (dropped it in 2018) and if I still did I'd subscribe to Geforce Now or just do console gaming.
Who says Mac users aren't gamers? It's a self perpetuating vicious cycle: gamers use Windows because that's where the majority of games are, so developers keep targeting Windows, so gamers use Windows, and so ad infinitum.
But people who use Macs enjoy games as well. They would rather not dual boot to Windows.
It's not true that Mac users don't play games. Rather, it's that most games are on Windows, which is a shame.
What I think you mean is competitive gamers and AAA gamers do not play on Apple hardware. This is mostly true today ofc, but keep in mind that's actually not the majority of the gaming market. Apple is raking it in from its gaming market.
I have no idea what the answer is. Personally I have run a number of games on macOS, via Steam, and via Boot Camp and virtualization. Some popular MMOs like Final Fantasy XIV, World of Warcraft, and Eve Online have macOS clients, though Guild Wars 2 discontinued theirs.
Apple Arcade apparently has enough Mac users that Apple has a reason to support it on macOS as well as iOS.
And Apple apparently has some reason to support iOS/iPadOS games on Apple Silicon Macs as well (though it could just be a side effect of a future iOS-macOS merger or hybrid device.)
...and how many would play more games on their Macs if they were available?
I refuse to believe Mac users are less fond of videogames. Based on my personal observation of the Mac users I know, they enjoy games as much as anyone.
At least casual games on iTunes for the iPad have a vast library, with many genuine good games (source: me, my iPad 2 was mainly a gaming platform for me, never found other uses for it).
I do realize more complex games are a different beast. But casual? I'd say Apple fans love them.
Can we stop it with the meme that these GPUs are unobtainable? Yes, they are still overpriced compared to their supposed original prices and they'll likely never return to that price given that the base prices of manufacturing, materials and such have increased for multiple reasons.
But stock has been generally available for many months now and it's possible to get them as long as you can afford them.
"M1 Ultra has a 64-core GPU, delivering faster performance than the highest-end PC GPU available, while using 200 fewer watts of power."
Note that they say "faster performance" not "more performance". What does "faster" mean? Who knows!
I still think “faster performance” sounds sound odd, but I understand their point.
My guess is they're putting a lot of weight on the phrase "relative power" and hoping you assume it means "relative to each other" and not "relative to their previous generation" (i.e. M1 Ultra -> M1 Max and RTX 3090 -> RTX 2080Ti) or "relative to the stock power profile".
Put bluntly, if the M1 Ultra was capable of achieving performance parity with an RTX 3090 for any GPU-style benchmark then Nvidia (who are experts in making GPUs) would have captured this additional performance. Bear in mind the claim seems to be (on the surface) that the M1 Ultra is achieving with 64 GPU cores and 800GB/s memory bandwidth what the RTX 3090 is achieving with 10,496 GPU cores and 936.2GB/s memory bandwidth.
But you'll all but certainly see the 3090 win more benchmarks (and by a landslide) than the M1 Ultra does. Because Nvidia is really, really fucking good at this, and they spend an absurd amount of money working with external projects to fix their stuff. Like contributing a TensorFlow backend for CUDA. Or tons of optimizations in the driver to handle game-specific issues.
Meanwhile Apple is mostly in the camp of "well we built Metal2, what's taking ya'll so long to port it to our tiny marketshare platform that historically had terrible GPU drivers?"
It is also worth noting that the M1 Ultra is an SoC so it'll have more than just CPU/GPU on it, by the looks of things it has some hefty amounts of cache, it'll also have a few IP blocks like a PCIe controller, memory controller, SSD controller (the current "SSDs" look to just be raw storage modules).
All told it likely still has somewhere in the region of 30-40 billion transistors for the GPU. Each GPU core being physically bigger than the 3090 is probably pretty good for some workflows and not so good for others. Generally GPUs benefit from having a huge number of tiny cores for processing in parallel, rather than a small number of massive cores.
Current benchmarks put it at roughly the performance of an RTX 3070, which is good for its power consumption, but not even close to the 3090. As I mentioned in the previous post, it just doesn't have the cores or memory bandwidth needed for the types of workloads that GPUs are built for (although unified memory being physically closer can help here ofc.), certainly not enough to make it a competitor for something like a 3090.
Edit: Oh also, for massively parallel workloads (like what GPUs do), more cores and better bandwidth to feed those cores will be one of the biggest performance drivers. You can get more performance by making those cores bigger (and therefore faster) but you need to crank the transistor count up a _lot_ to match the kinds of throughput that many tiny cores can do.
I really hope I'm wrong (as someone who owns an M1 Pro chip) but I find it hard to imagine things changing significantly in the next ~2 years unless someone is able (legally and technically) to release a CUDA compatibility layer.
Naturally, the HIP tooling doesn't support M1 GPUs at this time. We'll see if anyone else tries.
The GPU claims wouldnt even need to be on parity with NVIDIA, it would just need to offer a vertically integrated alternative to having to use EC2.
What reliability issues are you having with TensorFlow on M1 Macs?
Now i've got a team of data scientists in a fully MBP shop and we're holding off upgrades to M1 until this all gets resolved.
On my personal M1, I managed to make it work, but its hard to know the layers of changes made and what exactly allowed it to work.
You can buy single tensor accelerators from Google: https://www.coral.ai/products/
You can buy a bunch of those integrated into a single PCI-E card. https://iot.asus.com/products/AI-accelerator/AI-Accelerator-...
Cheap too. Some of these work with Mac. More of them work for PC, because the hardware interface is outside of Apple's thin vertical slice/garden.
Currently, you need this kind of GPU performance for high resolution VR gaming at 90 fps, but its just barely enough. This means that the GPU will run very loudly and heat up the room, and running games like HL Alyx on max settings is still not possible.
It seems that Apple might be the only company who can deliver a proper VR experience. I can't wait to see what they've been cooking up.
I could definitely see the Max and Ultra with a beefier cooling system (like the Studio’s) having pretty respectable performance.
M1 Max struggles to keep up with an RTX 3060 mobile.
Now that's with the overhead of Rosetta 2 and all that, so it's of course "not fair" for the M1. But that's also the current reality of the market, so ya know.
So publishers/developers need to make more native games. Even though every Mac port will probably make 1/10th revenue of a Windows title I guess Mac users would be happy to pay more for better games. I certainly would.
The Studio is a more compact form factor than any modern 4K gaming console. If they chose to ship something in that form factor with tvOS, HDMI, and an M1 Max/Ultra, it would be a very competitive console on the market — if game developers could be persuaded to implement for it.
How would it compare to the Xbox Series X and PS5? That’s a comparison I expect to see someday at WWDC, once they’re ready. And once a game is ported to Metal on any Apple silicon OS, it’s a simple exercise to port it to all the rest; macOS, tvOS, ipadOS, and (someday, presumably) vrOS.
Is today’s announcement enough to compel large developers like EA and Bungie to port their games to Metal? I don’t know. But Apple has two advantage with their hardware that Windows can’t counter: the ability to boot into a signed/sealed OS (including macOS!), load a signed/sealed app, attest this cryptographically to a server, and lock out other programs from reading with a game’s memory or display. This would end software-only online cheating in a way that PCs can’t compete with today. This would also reduce the number of GPUs necessary to support to one, Apple Metal 2, which drastically decreases the complexity of testing and deployment of game code.
I look forward to Apple deciding to play ball with gaming someday.
They could always choose to remedy that with a generous buyout offer.
Also, might be cheaper a couple years down the line.
They also don't need the display, camera, microphone. And could sell it at a loss and make the margins with TV+ and game sales.
But they would need their own bundled controller/accessories and get serious about AAA gaming.
Edit: https://arstechnica.com/information-technology/2022/01/pluto... says
> Microsoft already used Pluton to secure Xbox Ones and Azure Sphere microcontrollers against attacks that involve people with physical access opening device cases and performing hardware hacks that bypass security protections. Such hacks are usually carried out by device owners who want to run unauthorized games or programs for cheating.
So initially you could have Pluton-only servers and down the line non-Pluton hardware will simply be obsolete.
They won't have the Ultra GPU, but Apple's been shipping for years and Microsoft is just now bringing Pluton to market. I do wish them luck, but that's a lot of PC gamer hardware to depreciate.
I wish.
But playing ball is more than hardware. It is spending billions to buy Activision or Bungie. And I can't honestly imagine Apple having the cultural DNA or leader aspiration to make a game like The Last of Us where the player is brutally beating zombies to bloody clumps.
In video games the business side demands having exclusives, or timed exclusives, to sponsor twitch streamers playing your game and cutting special deals with studios. This is very different to the App store where Apple emphasizes their role as a neutral arbiter and a dev having the same deal as any other dev. Can you imagine the complains here on hn if Epic Games would get a a special deal just because they are a bigger fish and Fortnite is popular?
In the PC space the average spend is $800 and a PS5 is $500.
Yet the iOS 'gaming' scene, despite being one of the major revenue drivers, consists mostly of low-quality F2P games.
In any case, it's a moot point. Apple clearly doesn't care about desktop gaming and it shows in both their hardware and software.
Though every video game company on the planet hates them because of App Store terms.
I remember people saying this about phones in 2006.
This criticism is something that is a positive to me. Opposing companies are often dependant on adverting money and the things this leads to are a whole lot worse in my view.
Apple has invested billions into their gaming division. The big thing they need right now is a new version of Metal that gets feature parity with Vulkan or DX.
Also of note, there are very persistent rumors of an upcoming VR headset. Their M1 alone would blow away competition like the Quest. A Pro or max chip with some disabled CPU cores wouldn't cost a ton due to being scavenged cores and would positively stomp the competition.
Proton on ARM Mac's would involve Rosetta and while that does a surprisingly good job of running x86 on ARM I'm not sure it's up to the job of running games at high speed...
Some games are a stable 4k120 and others are more like 4k75.
I feel the 3090 could feasibly drive 5k60 as a result.
GTAV on Win11 Arm VM - okay, but not great.
GTAV on Crossover - much much better, lack of joystick support (but that's a crossover issue)
The advancements in such a short period of time in the amount of computing power, low power usage, size, and heat usage of these chips is unbelievable and game-changing.
This was probably further fueled by their soured relationship with Intel which was responsible for thermal issues on MBP's for years, poor performance increases across their entire Mac lineup, and poor cellular radio performance on some iPhone models -- forcing them to settle with Qualcomm and ditch Intel for mobile radios.
The performance of x86 has been a leader for a while mostly because of the sheer amount of optimization work that has gone into them, but the cruft of the x86 instruction set and the architectural stuff you have to do to make the instruction set work is really showing it's age.
That being said, the GPU performance claims are incredibly misleading. The previous "relative performance" benchmarks that were done on the M1 Max for GPU performance were misleading as well, they definitely cannot keep up with a mid-tier modern discrete GPU.
The GPU claim isn't an ARM/x86 comparison like the CPU performance would be. This is comparing a 64 core 800GB/s GPU with a 10k core 900GB/s GPU and trying to make them look equivalent through misleading marketing.
None of this is to say that the M1 Ultra is bad necessarily, even if it performs roughly the same as a mobile GPU or powerful iGPU it would still be a very good chip, and I'd love to use one if I could use it in my environment properly. I'm just saying don't put too much faith in the GPU performance measurements provided here.
Without denying some good work and engineering having gone into some x86 chips, they are not the reasons for the x86 becoming a leader. The duopoly of Intel and Microsoft – coupled with the aggressive Intel strategy to undermine competitors on the pricing and the sheer production volume they could quickly ramp up – has squeezed every single other viable competitor out of the market and relegated very few to become niche players (e.g. POWER) and entrenched the duopoly as a unfortunate leader. And then complacency and arrogance had set in for years to come until recently.
At least in the US You can get a 3090 for 2200 even with markup and an almost as good 3080 for $1150. If you wait until its in stock at a big box store you can get one for less even. A machine could be built with a 3080 for $2000.
Meanwhile a system based on on the 64 core GPU will run you $5000 and as such is affordable to nearly nobody thus few will get a chance to see it drastically under perform in the gaming arena on any of the games that don't support mac on arm.
With an absolutely invisible market share in the gaming desktop there will never be any incentive for anyone to change this insofar as direct support leaving you reliant on translation from x86 and from windows executables paying doubly in terms of compatibility and performance from an already very expensive and lackluster starting point.
Apple's footnotes don't even pretend to explain what these charts are.
> Performance was measured using select industry‑standard benchmarks.
The M1 announcement. They turned out to be pretty accurate.
So I’ll wait and see real benchmarks, but it wouldn’t surprise me if this does have incredible performance.
My understanding was that it was a reasonable comparison in some benchmarks, though when plugged in to power the mobile 3070 still had more headroom.
The problem was the implication that you'd get 3070 gaming performance. That was never going to be true because of the un-optimisation tax for games on Mac.
There doesn't exist a AAA game built for Metal and the Mac. The closest are games like World of Warcraft, Divinity Original Sin 2 – and even they are just "good ports" not originally designed for Mac (and are far from AAA graphics). This is why on Intel Macs, games under Bootcamp always ran 30%-50% faster, even though the hardware was the same.
Games on M1 Max run as you'd expect – about 30% slower than a 3070 for the same old reasons (and some new ones, like not being compiled for Apple Silicon at all). The GPU is about the same speed as a 3070 and it's doing what you'd expect, given the 30% unoptimization-tax workload.
- the workloads in games can vary a lot, vertex/fragment shaders imbalance, parallel compute pipelines, mixed precision (which the M1 gpu does not do), .. So another explanation is that you can get some 3070 parity on a cherry picked game, like a broken clock is right twice a day, but that does not make it generally true. Objective benchmarks have put the M1 gpus way slower than 3070 on average, and software support seems like an easy but false distraction given the Proton tax on Linux (which is not 30/50%)
- the M1 gpus are lacking a ton of hardware, matrix mul, fp16 again, ray tracing, VRR probably (not sure about this last one). These are used by modern games or applications, you may find a benchmark which skip them, but in the grand scheme of things it's something that the M1 gpu will have to emulate more often than not, and this has a cost
Waving all that as "the GPU is about the same speed" is technically wrong, or not really backed by facts at the very least
A lot of them, particularly the GPU benchmarks, were misleading because they only looked at performance that they had dedicated silicon for.
I run VR games on the index at 144hz with high settings without issue on a 3060 Ti.
I've been on the market for a 3070-3090, but only because I want a card for which a water block is available, not because I need more power for any extant game.
Really looking forward to Apple's VR offering after seeing the performance of their compact SoC
Not to say it'll never happen, but its not a done deal basically, and to my knowledge the process hasn't yet started
So expect what that graph actually means is some extremely specific, cherry-picked benchmarks
[0]https://www.amd.com/en/products/cpu/amd-ryzen-threadripper-3...
There's no other chip that has the power of an RTX 3090 and more power than an i9-12900K in it - after all, Threadripper doesn't have a lick of graphics power at all. This chip can do 18 8K video streams at once, which Threadripper would get demolished at.
I'm content with giving them the chip crown. Full system? Debatable.
Edit: even AMD themselves call their threadripper lineup workstation chips, not personal.
If the purchase page says to "contact sales" and doesn't list a price then it is not for consumers.
I mean they all are CPUs coming out this year as far as I know.
But you know, I'm still happy to double my current build perf in a small box I can stick in my closet. Ordered one :-)
Also 1st gen threadrippers are getting on a bit now, surely. It's a ~6 year old microarchitecture.
I am curious whether there is a real performance difference?
I do lots of computing on high-end workstations. Intel builds used to be extremely expensive if you required ECC. They used that to discriminate prices. Recent AMD offerings helped enormously. I wonder whether these M1 offerings are a significant improvement in terms of performance, making it worthwhile to cope with the hassle of switching architectures?
The above statement should also relate to most other C/C++ projects.
Threadripper 3990X get about 25k in Geekbench Multicore [1]
M1 Max gets about 12.5k in Geekbench Multicore, so pretty much exactly half [2]
Obviously different tasks will have _vastly_ different performance profiles. For example it's likely that the M1 Ultra will blow the Threadripper out of the water for video stuff, whereas Threadripper is likely to win certain types of compiling.
There's also the upcoming 5995WX which will be even faster: [3]
[1] https://browser.geekbench.com/processors/amd-ryzen-threadrip...
[2] https://browser.geekbench.com/v5/cpu/search?utf8=%E2%9C%93&q...
[3] https://www.amd.com/en/products/cpu/amd-ryzen-threadripper-p...
Maybe the cooling and power delivery difference between laptop formfactors and PC formfactors will be less with these new arm based chips.
If I was to guess, the increased cooling probably helps the Studio sustain similar boost clocks as the laptops, but for longer.
Although it’s possible these are on N4x, which might increase the attainable boost.
24-core scores 20k, 32-core scores 22.3k, and 64-core score 25k. Something isn't scaling there.
It sounds pointless to come up with synthetic benchmarks which emulate software that is not able to handle hardware, and then use said synthetic benchmarks to evaluate the hardware performance.
Most consumers are software aware, not hardware aware. They care what they will use the hardware for, not what they can use it for. To that end, benchmarks that correlate with their experience are more useful than a tuned BLAS implementation.
I use a few 32 and 64 core machines for build servers and file servers, and while the 64-core EPYCs are not twice as fast as the 32-core ones due to lower overall frequency, they're 70% or so faster in most of the things I throw at them.
I was under the impression that all of their multi-core tests were "run N independent copies of the single-threaded test", just like SPECrate does.
At the same time the Threadrippers also have a gargantuan amount of cache that can be accessed at several hundred gigabytes per second per core. Obviously not as nice as being able to hit DRAM at that speed.
Also, not everything fits into cache.
You could still write OpenCL kernels of course. Doesn't mean you can't use it, but not sure if it's all just accessible to CPU-side code.
(or maybe it is? it's still a damn fast piece of hardware either way)
Linking this[1] because TIL that the memory bandwidth number is more about the SoC as a whole. The discussion in the article is interesting because they are actively trying to saturate the memory bandwidth. Maybe the huge bandwidth is a relevant factor for the real-world uses of a machine called "Studio" that retails for over $3,000, but not as much for people running postgres?
1 - https://www.anandtech.com/show/17024/apple-m1-max-performanc...
https://semiaccurate.com/2022/03/08/amd-finally-launches-thr...
The niche for high clocks was arguable with the 2nd-gen products but now you are foregoing v-cache which also improves per-thread performance, so Epyc is relatively speaking even more attractive. And if you take Threadripper you have artificial memory limits, half the memory channels, half the PCIe lanes, etc, plus in some cases it's more expensive than the Epyc chips. It is a lot to pay (not just in cash) just for higher clocks that your 64C workloads probably don't even care about.
AMD moved into rent-seeking mode even before Zen3 came out. Zen2 threadripper clearly beats anything Intel can muster in the segment (unless they wanted to do W-3175X seriously and not as a limited-release thing with $2000 motherboards) and thus AMD had no reason to actually update this segment when they could just coast. Even with this release, they are not refreshing the "mainstream" TRX40 platform but only a limited release for the OEM-only WRX80 platform.
It was obvious when they forced a socket change, and then cranked all the Threadripper 3000 prices (some even to higher-levels than single-socket Epyc "P" skus) what direction things were headed. They have to stay competitive in server, so those prices are aggressive, but Intel doesn't have anything to compete with Threadripper so AMD will coast and raise prices.
And while Milan-X isn't cheap - I doubt these WRX80 chips are going to be cheap either, it would be unsurprising if they're back in the position of Threadripper being more expensive for a chip that's locked-down and cut-down. And being OEM-only you can't shop around or build it yourself, it's take it or leave it.
The performance per watt isn’t in the same universe and that matters.
I couldn’t give less of a shit about performance-per-watt. The ONLY metric I care about is performance-per-dollar.
A Mac Studio and Threadripper are both boxes that sit on/under my desk. I don’t work from a laptop. I don’t care about energy usage. I even don’t really care about noise. My Threadripper is fine. I would not trade less power for less noise.
One hour of my time is more expensive than an entire month of a computer electricity bill.
Some people just want tasks to perform as fast as possible regardless of power consumption or portability.
Life's short and time is finite.
Every second adds up for repetitive tasks.
Power does make a big difference in data centers though - it's often the case that you run out of power before you run out of rack space.
Where power for a computer might make a difference could be in power-constrained (solar/off grid) scenarios.
I don't know if I've ever heard anyone make an argument based on $$$.
I personally stick to the lower wattage ones because I don't generally need high end stuff, so I think Apple is going the right direction here, but it should be noted that Intel has also started down the path of high performance and efficiency cores already. AMD will find itself there too if it turns out that for home use, we just don't need a ton of cores, but instead a small group of fast cores surrounded by a bunch of specialist cores.
Thermal density plays a huge role, the size of the chips is going down faster than the wattage, so thermal density is going up every generation even if you keep the same number of transistors. And everyone is still putting more transistors on their chips as they shrink.
Going forward this is only going to get more complicated - I am very interested to see how the 5800X3D does in terms of thermals with a cache die over the top of the CCD (compute die). But anyway that style of thing seem to be the future - NVIDIA is also rumored to be using a cache die over the top of their Ada/Lovelace architecture. And obviously 60W direct to the IHS is easier to cool than 60W that has to be pulled through a cache die in the middle.
Looking it up though I do see a lot of concerns with the heat they generate. I can only conclude I don't push my chip very hard (which, honestly, I probably don't)
I've been happy with the AMDs I purchased over the past 4 years, we'll see how they hold up and how this next gen comes out. I did see that the recent Intels are quite competitive which is good for everybody.
Yeah, longevity, blah blah, but laptop chips are designed to sit above 90C under load, it's fine.
Just saying that "how hard it is to cool" doesn't solely depend on power consumption anymore. Heat density is making that harder and harder, even if power consumption stays the same.
What does improve though is how much heat it pumps into your room. Yeah, a Rocket Lake at 200W might be roughly as hard to cool as an AMD at 90W or whatever... but one is still putting 200W into your room and the other is still putting 90W. Temperatures are not the same thing as power dissipation either. I don't like having my gaming PC running in my room during the summer, and I'm actually looking at maybe running cables through the walls to have it in the basement instead. I also have a 5700G and some NUCs that are much lower power that I prefer to use for surfing and shitposting.
Sure it does. Reading the rest of your post I think you're more talking about temperature than cooling requirements, but a 200W CPU needs 200W of heat dissipation, while a 60W CPU only needs 60W of heat dissipation. It's literally a 1:1 relationship since CPUs don't do any mechanical work, so power in == heat out.
Keeping temperatures below some arbitrary number does then include things like density, IHS design, etc... But that only matters for something like Intel's "Thermal Velocity Boost" where it's really important to stay under 70C specifically instead of just avoiding thermal throttling.
Until maybe these M1's (and I'm not entirely convinced) I've not in the 20 years I've been computing seen a reasonably configured desktop (eg not just a laptop on a stick ala iMac but an ACTUAL desktop) ever not smoke the pants off of every single laptop you could put up against it. It's hard to beat the one-two punch of lots of power and room to cool it. If you are sitting at at desk why the heck wouldn't you leverage that?
I still have a proper desk-based working environment hooked up to a docking station though. I really wouldn't want to use a laptop that doesn't have a first-party dock as my primary machine.
I agree that most developers are web/mobile developers who use a laptop. That’s great. I am an increasingly niche developer.
The root comment was a comparison against Threadripper. Normal developers should not waste money on a Threadripper. If someone is a niche developer that warrants a Threadripper then pointing out that most developers don’t need a Threadripper is a waste of time.
https://www.theguardian.com/society/2021/sep/09/transport-no...
As for desktops, watercooling makes computers dead silent.
I personally bought a Ryzen 5950(?) instead of a Threadripper because I figured I'd accidentally spill water all over it or however it works. There are not many watercooled OEM products as far as I know.
The M1 Ultra is a workstation part. It goes in machines that start at $4,000. The competition is Xeons, Epycs, and Threadrippers.
Our "world" build is slightly faster on my M1 Max.
https://twitter.com/kiratpandya/status/1457438725680480257
The 3990x runs a bit faster on the initial compile stage but the linking is single threaded and the M1 Max catches up at that point. I expect the M1 Ultra to crush the 3990x on compile time.
Curiosity got the better of me:
Isn't linking IO-bound?
https://llvm.org/devmtg/2017-10/slides/Ueyama-lld.pdf
There is a breakdown in those slides discussing what parts of lld are single threaded and hard to parallelize so I suspect single thread performance plays a big role too. I generally observe one core pegged during linking.
That would mean that these comparisons between Threadripper and the M1 Ultra do not reflect CPU performance but instead showcase whatever choice of SSD they've been using.
Why did you omit the reference to "file system"?
Are we supposed to ignore the fact that a linker's main job is reading object files and write the output to a file?
I find this sort of argument particularly comical given a very old school technique to speed up compilation is to use a RAM drive to store the build's output.
Does it, though?
I mean, if you read that link you'll notice it boasts the linker's performance by comparing it with cp and how it's "so fast that it is only 2x slower than cp on the same machine."
Is cp supposed to be CPU-bound?
Just that with the same hot caches, the average change-build-test loop that developers do 100+ times a day is just faster on the M1 Max.
(+ now I see it's rust: how parallel is your build, really?)
Not the OP but I install a lot of Rust projects with Cargo and recently did some benchmarking on DigitalOcean's compute-optimized VMs. Going from 8 cores to 32 cores was a little disappointing:
Bat (~40 crates): 68s -> 61s
Nushell (486 crates): 157s -> 106s
Compilation starts out highly parallel and then quickly drops down to a small number of cores.
If the final x86 production build takes longer it doesn’t matter - that happens on the cloud anyway.
Edit: Rust builds are very parallel until linking. No different than any other LLVM build.
It matters when comparing CPU performance, which is what this benchmark is being used for.
Try the same thing with mold.
So, even if it doesn't quite beat Threadripper in the CPU department - it will absolutely annihilate Threadripper in anything graphics-related.
For this reason, I don't actually have a problem with Apple calling it the fastest. Yes, Threadripper might be marginally faster in real-world work that uses the CPU, but other tasks like video editing, graphics, it won't be anywhere near close.
We all need to take Apple claims with grain of salt as they are always cherrypicked so i wont be surprise if it wont be even 3070 performance in real usage.
Don't worry though there will still be room to move the goalposts with "uhhh, but, Apple is designing for high IPC and low clocks, it's totally different and x86 could do it if they wanted to but, uhhh, they don't!".
(I'm personally of the somewhat-controversial opinion that x86 can't really be scaled in the same super-wide-core/super-deep-reorder-buffer fashion that ARM opens up and the IPC gap will persist as a result. The gap is very wide, higher than 3x in floating-point benchmarks, it isn't something that's going to be easy to close.)
Work out the IPC there - the Intel has a 2x thread count advantage, a 17% clock advantage, and Apple comes out 5% ahead. So the IPC gap there is about 2.46x.
It's not a perfect comparison of course, since we're mixing SMT and big/little cores, but in basically every area Intel should (on paper) have more resources available and Apple is coming out on top anyway by sheer IPC.
That's what I'm saying - you can't really do that approach with x86. It's not power-advantageous or transistor-advantageous to go super wide on the decode or reorder buffer like that on x86. And regardless of the tricks x86 uses to mitigate it, you've still got a 2.5x IPC gap at the end of the day. A 2.5x IPC gap will not be closed up by just a single node shrink.
And that's looking at MT, where your task scales perfectly. See where I'm going with this? Intel is using 2x the number of threads, and 3x the number of efficiency cores to get there. Apple can deliver that punch across a much lower number of threads - meaning ST-bottlenecked tasks will scale much much better on Apple.
With a single-threaded test, the M1 is pulling 7W vs 33W for the Alder Lake intel. Obviously that tells us nothing about efficiency, since we'd need to know the scores, but that's the downside, is for normal, poorly-threaded tasks, like surfing the web or editing code, the 12900HK is going to be boosting high to reach the same performance levels the M1 does at 3 GHz. And that's exactly what you see in the power figures there.
In short: you will likely see x86 able to keep up in one metric or another. You can win on performance if you just go nuclear on power. You can match on power on perfectly-threadable tasks that allow the x86 to deploy twice the threads (sharing instruction cache/etc). You can match on single-threaded battery life if you accept lesser performance. But the overall performance of the M1 derives from the massive IPC it generates, and that's something that x86 can't match nearly as easily.
Going ham on a single metric just to claim victory isn't nearly the same thing as the level of all-round performance and efficiency that Apple has achieved there.
(see also, putting a 128-thread Threadripper 3990WX workstation up against a 10-thread M1 Max laptop just to win at rendering... and people here thought that disproved that Apple was great hardware lol)
The reorder buffer size is just a logical consequence of the frontend width.
And yes, scaling an Aarch64 frontend is dead simple compared to x86 due to the fixed instruction width. The disadvantage of x86 is serious, but I don't know if we can count it out quite yet. This is the first time Intel and AMD got any serious pressure on that front. I'm sure they're taking the challenge seriously, and it'll take some years before we'll see the results.
And since the presenter mentioned the Mac Pro would come on another day, I wonder if they'll just do 4x M1 Max for that.
Well they were only correct that Apple managed to hide a whole section of Die Image. ( Which is actually genius ) Otherwise it wouldn't have made any sense.
Likely to be using CoWoS from TSMC [1] since the bandwidth numbers fits. But needs further confirmation.
I wrote about it three months ago.
They'll be running out of names for that thing. M1 Ultra II would be lame, so M1 Extreme? M1 Steve?
There might be some hesitance installing an M5. You should stay out of the way if the machine learning core needs more power.
I guess by the time they get to M5, anyone old enough to get the reference will have retired.
That would be the funniest thing Apple has done in years. I totally support the idea.
Elon is somewhat toxic these days...
The thing is they're at 128GB which is way way far from 1.5TB. You're not going to find a way to get 12x the memory while still doing the embedded memory packages.
Maybe I'll be pleasantly surprised but it seems like they're either going to switch to (R/LR)DIMMs for the Mac Pro or else it's going to be a "down" generation. And to be fair that's fine, they'll be making Intel Mac Pros for a while longer (just like with the other product segments), they don't have to have every single metric be better, they can put out something that only does 256GB or 512GB or whatever and that would be fine for a lot of people.
https://www.anandtech.com/show/17058/samsung-announces-lpddr...
> It’s also possible to allow for 64GB memory modules of a single package, which would correspond to 32 dies.
It is possible, and I guess that NVIDIA’s Grace server CPU will use those massive capacity LPDDR5X modules too.
The M1 Ultra has 8 memory packages today, and Apple could also use 32-bit wide ones (instead of 64-bit) if they want more chips.
You (or the OS or the chip) could page things in and out if the unified memory. Treat unified memory as a MEGA L3 cache.
Depending on how it’s done it may not be transparent if you want the best performance. But would it work?
Unlikely, M1 Ultra is the last chip in the M1 family according to Apple [1].
"M1 Ultra completes the M1 family as the world’s most powerful and capable chip for a personal computer.”"
[1] https://www.apple.com/newsroom/2022/03/apple-unveils-m1-ultr...
> The Mac is now one year into its two-year transition to Apple silicon, and M1 Pro and M1 Max represent another huge step forward. These are the most powerful and capable chips Apple has ever created, and together with M1, they form a family of chips that lead the industry in performance, custom technologies, and power efficiency.
I think it is just as likely that they mean "completes the family [as it stands today]" as they do "completes the family [permanently]."
[1] https://www.apple.com/newsroom/2021/10/introducing-m1-pro-an...
edit: This comment around SoC code names is worth a look too: https://news.ycombinator.com/item?id=30605713
They’d need that anyway for a Mac Pro replacement (128GB wouldn’t cut it for everyone), but even for smaller config it’s frustrating being limited to 16G on the M1 and 32 on the Pro. Just because I need more RAM doesn’t mean I want the extra size and heat or whatever.
Since I run a lot of memory intensive tasks but few CPU or GPU bound tasks, a regular m1 with way more memory would be ideal.
Judging from the geekbench scores[0], M1, M1 Pro, and M1 Max perform identically in single threaded tasks. And the newly leaked Mac Studio benchmark[1] shows essentially identical single thread performance.
[0]: https://browser.geekbench.com/mac-benchmarks [1]: https://browser.geekbench.com/v5/cpu/13330272
Turns out I was right.
The Mac Pro chip will be a different thing/die.
Ehrm, anyway.
It actually isn't clear to me whether designing a two socket motherboard is fundamentally an easier task than jamming more of the things into a single package (given that they have already embraced some sort of chiplette paradigm).
or they could take a page out of microsofts book and just call the next one "m one"
(edit: per a sibling comment, if the internals like IRQ only really scale to 2 chiplets that pretty much would rule it out though.)
Could AMD/Intel follow suit and package memory as an additional layer of cache? I worry that we are being dazzled by the performance at the cost of more integration and less freedom.
edit: typo + stacking + rumoured date
It's at least two separate chips combined together. That makes more sense, mitigates the problem.
There's already Reddit if you want to crack puns and farm karma. Let's try to keep the signal:noise ratio higher here.
On the other hand I wonder what exactly it can do. To what degree are you tied into a specific neural architecture (eg recurrent vs convolutional), what APIs are available for training it, if it's even meant to be used that way (not just by Apple-provided featues lke FaceID)?
https://developer.apple.com/machine-learning/
https://developer.apple.com/documentation/coreml/model_custo...
----
Mac Pro scale up?
How is this going to scale up to a Mac Pro, especially related to RAM?
The Ultra caps at 128 GB of RAM (which isn't much for video editing, especially given that the GPU uses the system RAM). Today's Mac Pro goes up to 1.5TB (and has dedicated video RAM above this).
If the Mac Pro is say, 4 Ultra's stacked together - that means the new Mac Pro will be capped at 512GB of RAM. Would Apple stack 12 Ultra's together to get to 1.5TB of RAM? Seems unlikely.
Then, on the highest configuration, I think they actually can put 6 M2-top-specced or more into the Mac Pro.
But yes, I see a lot of folks replacing current Mac Pros with Studios.
- the shared CPU+GPU RAM doesn't necessarily mean the GPU has to eat up system RAM when in use, because it can share addressing. So whereas the current Mac pro would require two copies of data (CPU+GPU) the new Mac studio can have one. Theoretically.
- they do have very significant video decoder blocks. That means that you may use less RAM than without since you can keep frames compressed in flight
I'd expect it to work more like a game console, streaming in content from the SSD to working memory on the fly, processing it with the CPU and video decode blocks, and insta-sharing it with the GPU via common address space.
All that is to say, where you needed 1.5TB of RAM on a Xeon, the architectural changes on Apple Silicon likely mean you can get away with far less and still wind up performing better.
The "GHz myth" is dead, long live the "GB myth."
The RAM is not on die. It’s just soldered on top of the SoC package.
> All that is to say, where you needed 1.5TB of RAM on a Xeon, the architectural changes on Apple Silicon likely mean you can get away with far less and still wind up performing better.
No, it does not. You might save a bit, but most of what you save is the transfers, because moving data from the CPU to the GPU is just sending a pointer over through the graphics API, instead of needing to actually copy the data over to the GPU’s memory. In the latter case, unless you still need it afterwards you can then drop the buffer from the CPU.
You do have some gains as you move buffer ownership back and forth instead of needing a copy in each physical memory, but if you needed 1.5TB physical before… you won’t really need much less after. You’ll probably save a fraction, possibly even a large one, but not “2/3rd” large, that’s just not sensible.
They just can't ship a Mac Pro without expansion in the normal sense, my guess is that the M2 will combine the unified memory architecture with expansion busses.
Which sounds gnarly, and I don't blame them for punting on that for the first generation of M class processors.
https://en.wikipedia.org/wiki/List_of_Apple_codenames
M1 Max is Jade C-Die => 64GB
M1 Ultra is Jade 2C-Die => 128GB
There is a still unreleased SoC called Jade 4C-Die =>256GB
So I think that's the most we'll see this generation, unless they somehow add (much slower) slotted RAM
If they were to double the max RAM on M2 Pro/Max (Rhodes Chop / Rhodes 1C), which doesn't seem unreasonable, that would mean 512GB RAM on the 4C-Die version, which would be enough for _most_ Mac Pro users.
Perhaps Apple is thinking that anyone who needs more than half a Terabyte of RAM should just offload the work to some other computer somewhere else for the time being.
I do think it's a shame that in some ways the absolute high-end will be worse than before, but I also wonder how many 1.5TB Mac Pros they actually sold.
Whether they use slotted RAM or not has nothing to do with performance. It's a design choice. For the mobile processors it makes total sense to save space. But for the Mac pro they might as well use slotted RAM. Unless they go for HBM which does offer superior performance.
https://ark.intel.com/content/www/us/en/ark/products/134599/...
4800 MT/s is the actual maximum spec, anything beyond that is OC.
A16 would give great performance, and I think it’s safe for them to have a two year iteration time on laptop/desktops vs one year for phone/tablet.
It is strange Apple didn't cooperate with Intel in this area.
You’ll most likely also be able to buy dedicated GPUs/ML booster addon Cards and the likes for it.
It’s the most likely thing to happen or they won’t release another Mac Pro.
That said, the idea that USB C/Thunderbolt is the new PCIe bus has some merit. I have yet to find someone who makes a peripheral card cage that is fed by USBC/TB but there are of course standalone GPUs.
Oh please hell no.
I have to unplug and plug my USB-C camera at least once a day because it gets de-enumerated very randomly. Using the best cables I can get my hands on.
File transfers to/from USB-C hard drives suddenly stop mid-transfer and corrupt the file system.
Don't ask me why, I'm just reporting my experiences, this is the reality of my life that UX researchers don't see because they haven't sent me an e-mail and surveyed me.
Never had such problems with PCIe.
Sounds like you're listing the common complaints with usb-3 over usb-c peripherals, which are not a suitable replacement for PCIe. Thunderbolt is something different, more powerful & more reliable.
TBH, even a thumb drive would have me pissed if it disconnected at random times. That's what I hated about using the SD slot of MacBooks to host a semi-permanent drive.
So, I'm in total agreement.
Very interesting stuff. I wonder both if the Zynq Ultrascale RFSOC PCIe card would work in that chassis and if I could get register level access out of MacOS.
No need to run inside the kernel for these things any more.
I hope we get closer to that long-standing dream over the next few years.
But right now you can see laptop manufacturers so desperate to avoid thunderbolt bottlenecks that they make their own custom PCIe ports.
For the longest time, thunderbolt ports were artificially limited to less than 3 lanes of PCIe 3.0 bandwidth, and even now the max is 4 lanes.
Since when did the average developer care about how many sockets a mobo has...?
Surely you still have to carefully pin processes and reason about memory access patterns if you want maximum performance.
My understanding was the the dustbin was designed with one big processor because SMP/numa was a massive pain in the arse for the kernel devs at the time so it was easier to just drop it and not worry.
Or am I out of date on NUMA systems?
Remember the Pentium D? Unfortunately, I used to own one.
This stuff rapidly starts to make my head spin. I have not studied interconnects and have never written any NUMA-aware software. I will just post this link (read the "Memory Latency" section):
https://www.anandtech.com/show/16529/amd-epyc-milan-review/4
As I understand it, the I/O die is partitioned into four quadrants. Each quadrant has two memory controllers and is attached to two compute dies. CPUs can access memory attached to the same quadrant with lower latency than going to another quadrant. This is a NUMA system that can be configured to appear as one logical NUMA node.
I believe their smaller parts with two or fewer compute dies will be UMA, but with the same non-uniform latency to L3.
>So is that 56 core Xeon that Intel was bragging about for a while there until the 64 core Epycs & Threadrippers embarrassed the hell out of it.
I believe the 64-core Epycs and Threadrippers came first. The 56-core Xeon was a purpose-built part for HPC, so it wasn't quite a marketing gimmick.
Eg simulation softwares often used in the industry (but the one I’ve on top of my head is Windows only.)
Anyway, the point the make is this: if you claim doubling performance, but only the selected few softwares as you observed would be optimized to take advantage of this extra performance, then this is mostly useless to the average consumer. So their point is made exactly with your observation in mind, that all your softwares is benefiting from it.
But actually their statement is obviously wrong for people in the business—this is still NUMA and your software should be NUMA aware to be really squeezing the last bit of performance. It just degrades more gracefully to non optimized code.
This is a tragedy for the future of computing. It might as well be encased in resin. Great performance, but I won't spend car money on something I can't upgrade or repair.
It works for them because most of their products have only one or maybe two choices. It would never fly for white box sales, but Apple is not in that market.
It sounds like this operates as if it was one giant physical chip, not two separate processors that can talk very fast.
I can’t wait to see benchmarks.
It's probably best to think of this chip as an extremely fast double socket SMP where the two sockets have much lower latency than normal. Software written with that in mind or multiple programs operating fully independent of each other will be able to take massive advantage of this, but most parallel code written for single socket systems will experience reduced gains or even potential losses depending on their parallelism model.
Whereas AMD's solution is focused on increasing the cache size (hence the 3D stacking), Apple here seems to be connecting the 2 M1 Max chips more tightly. It's actually more reminiscent of AMD's Infinity Fabric interconnect architecture. https://en.wikichip.org/wiki/amd/infinity_fabric
The interesting part for this M1 Ultra is that Apple opted to connect 2 existing chips, rather than design a new one altogether. Very likely the reason is cost - this M1 Ultra will be a low volume part, as will be future iterations of it. The other approach would've been to design a motherboard that sockets 2 chips, which seems would've been cheaper/faster than this - albeit at expense of performance. But they've designed a new "socket" anyway due to this new chip's much bigger footprint.
Intels upcomming Saphire Rapid server CPUs are extremly similar, with wide connections between two close dies. Crossectional bandwith is in the same order of magnitude there.
I don’t see Apple getting back into that business. But I think they have the ability to make a good option if they want.
Intel is believed to have pretty good margins on their server CPUs
> and its outside their area of expertise.
That's what people used to say about Apple doing CPUs in-house.
Folks who said CPUs weren't their core expertise (I assume back in 2010 or before, prior to A4) missed out on just how involved they've historically been, what it takes to get involved, the role of fabs and off the shelf IP to gradually build expertise, and what benefits were possible when building silicon and software toward common purpose.
Lower TDP = lower electric bills and lower airconditioning bill. Win win
The M1 Ultra is already a little light on memory for its price and processing power; it would have much too little memory for a cloud host.
The server market is different. Companies buy servers from the low bidder. Apple has never really played in that market.
It will be interesting to see the difference in performance and performance per watt, when both companies are on the same node.
Besides for some use cases, these Mac Studios will be racked and in data centers as is.
Click & deploy from within Xcode (I hate Xcode though.)
Edit: Apparently in iPhones, they are used for FaceID.
Or, there will be some new form of video generation (like the ones generating video from Deep Dream etc, but something aimed at studio production) using ML that wasn't practically usable before.
It opens many doors, but it will take at least many months, if not years, to see some new "kind" of software to emerge that efficiently makes use of them.
https://www.digitalcameraworld.com/news/apple-m1-chip-makes-...
[1] https://developer.apple.com/documentation/accelerate/bnns
[2] https://www.tensorflow.org/lite/performance/coreml_delegate
Apple mentioned that the M1 Ultra is the last member of the M1 family. So how is the Mac Pro going to scale?
Will Apple enable “traditional” scaling by allowing multiple M1 Ultra chips to be combined in one system? Or what?
Further, how will an Apple Silicon based Mac Pro be made expandable; something that has been a corner stone feature of Mac Pros in the past?
Intel's got a lot of work to do to catch up. I think the only way Intel will catch up is to completely embrace RISC-V
/takes off foil hat
I wouldn’t expect single-threaded improvement until the M2.
When I get home, it's all about the GPU on my gaming PC (Windows). It's just that CPU just doesn't seem to be a huge bottleneck for me on the desktop anymore. Are Mac's different somehow where they need more CPU?
If you are doing CAD, things like fluid/particle physics simulations can really slam the CPU. The M1 Ultra isn't marketed to the normal user just doing some web browsing. Its the top tier chip for people who find the M1 insufficient.
Maybe not a huge caveat, as 16-core chips in the same power envelope probably covers most of what an average PC user is going to have, but there are 64-core Threadrippers out there available for a PC (putting aside that it's entirely possible to put a server motherboard and thus a server chip in a desktop PC case).
I'd like to see the actual performance comparison.
Alder Lake has been repeatedly shown to outperform M1 core-per-core. The M1 Ultra is just way bigger. (And way more power efficient, which is a tremendous achievement for laptops but irrelevant for desktops.)
Is it great leadership? Top tier engineering talent? Lots of money? I simply don't understand.
Hopefully leadership is really looking hard at this trend and adjusting future offerings accordingly. Consumers WANT machines with high performance and great I/O and they're willing to pay for them.
With Apple, Intel, and AMD really stepping up the last couple of years, I think the next decade of personal computing is going to be really exciting!
You have to remember that since the 2014 retina, Apple's offerings have been a bit crap.
This is a return to form ( and a good one at that) but its not worthy of hero worship. They've done a good job turning things around, which is very hard.
They have talent, they have execution, they have data about what rubs on macs they can use to optimize really well. But they have the profits and the cash reserves to make big bets and wait them out.
I think the M1 was expected 1 or 2 years before it was released. But they waited. Maybe it wasn’t good enough. Maybe the software support wasn’t there. But they didn’t have to push it out anyway and hope for the best. They could afford to wait.
Maybe that makes them willing to take bigger risks. Maybe through history they just knew Intel slowing down would happen (it bit them with 68k, then PPC, then G3/4/5) and we’re prepared in a way only done one with their own chips could be.
(edit: calm down people, I recognize it's impressive, but it's just not as fun an announcement as an architecture rev, which I was hoping for after a year :D )
2 : to draw something from as if by milking: such as
b : to draw or coerce profit or advantage from illicitly or to an extreme degree : exploit
milk the joke for all it's worth
https://www.merriam-webster.com/dictionary/milkThere are even some with 640GB. This is at a different price point though.
Rendering in VR takes a lot of memory at higher resolutions.
The board has been set. M1 Endgame is nearly ready.
Thinking is a superpower even better than being the first species to develop sight.
See also “The Last Question” by Asimov.
But your everyday apps like your browser have been fast for at least a decade.
Answer: Just wait for some genius to figure out how to run Electron inside Electron, and port Calculator to it.
The bottleneck for most slow/old PCs is not actually the CPU. Dried thermal paste, mismatched low speed RAM sticks, operating systems installed to mechanical disks and weak iGPUs are often to blame. You can fix most of these easily at home for very little money and keep using old systems.
Most people are sadly too tech-illiterate to really understand what an SSD is, so I still sometimes do SSD upgrades on almost brand new laptops and pre-built desktops for friends and family. OEMs like to add insult to injury by bloating up the default Windows install with janky driver control panels and adware. A new SSD, a clean OS install with generic drivers, new thermal paste and some compressed air in the fans will make those PCs work better than they did on day one.
I have a few of those kind of machines around the house...
But surely the GPU things can't be real? The GPU in the M1 Ultra beats the top-of-the-line Nvidia? That's nuts.
Can't wait for the (real world) reviews to be published
Hopefully AMD, Nvidia, others can follow the trend
Dubious. https://www.pcgamer.com/apple-m1-max-nvidia-rtx-3080-perform...
> Apple even says its new GPU is a match for Nvidia's RTX 3080 mobile chip, though you'll have to take Apple's word for it on that one. We've also reached out to Nvidia to see what it might have to say on the matter.
> RTX 3080 mobile chip
> mobile chip
There's a 50%[1] (!) difference with mobile and non-mobile versions of the chip. So that's hardly a deal breaker.
Even more incredible Anandtech reports the M1 max GPU block maxing at 43W in their testing. So a 90W GPU in the M1 Ultra is trading blows with a 350+ watt 3090.
1) https://www.anandtech.com/show/17024/apple-m1-max-performanc...
Might be nice for e.g. ML where you can effectively treat them as entirely independent GPUs but for games I would be surprised if this matches a high end GPU.
The first benchmarks are in, and it's mostly a 40-50% improvement in the M1 Max, not 100%.
In borderlands it got 24 FPS while the 3080 got 52 FPS. How is that on par?
Legacy games written for x86 CPUs obviously are going to perform poorly. I recommend you actually read the review and don't just scroll to the worst gaming benchmark you can find.
Maybe, but the "raw power" is useless if it can't be exploited.
> Legacy games written for x86 CPUs obviously are going to perform poorly.
Not if they're GPU-bound. Even native performance isn't that impressive
It’s a great chip but it doesn’t trade blows with anything Nvidia puts out especially at comparable price points.
Maybe you buy things to run benchmarks. I buy them to run the software I own. For games they come up short on fps and high on price. That is the inverse of what I’m looking for.
However if your workloads are in a more professional domain as mine are then it's entirely fair to say this chip is trading blows with Nvidia's best at lower prices. Don't forget this is an entire SOC and not just a GPU, power saving aren't irrelevant either if you actually work your hardware consistently as I do.
People that game on Mac know it's a lie, GPU for gaming on mac is vastly slower than recent graphic cards.
For some workloads i would not be surprised at all. But for all workloads, ...
We don't know yet. Apple is benchmarking against Workstation graphics cards
"production 2.5GHz 28-core Intel Xeon W-based Mac Pro systems with 384GB of RAM and AMD Radeon Pro W6900X graphics with 32GB of GDDR6"
From the linked article. Apple is comparing against RTX 3090.
For most non-parallel tasks, my guess is the Intel 12900K will beat at performance.
Intel's next generation will have 50% more cores and beat this chip at multithreading.
Nearly everything I use daily is built for M1 now.
https://isapplesiliconready.com/
And honestly, if it's not, its a good indication that it's time to move away from that product as they don't care about a huge segment of their users.
I want to see proper benchmarks before getting too exited.
[0] https://www.notebookcheck.net/The-new-MacBook-Pro-14-only-ma...
The pro and max come only in laptops, so the cooling difference should be quite significant, but also there is more chip and an interconnect to cool. Really looking forward to the in depth analysis of this.
Did they get their terminology confused? Later it says "By connecting two M1 Max die with our UltraFusion packaging architecture [...]" which also sounds like it's a MCM and not a SoC.
https://www.deseret.com/1999/9/1/19463524/apple-unveils-g4-d...
I kinda believe em this time, but time will tell.
Without having to be a kernel hacker, that is.
The now outdated Mac Pro goes up to 1.5TB, only 128GB available here.
Hopefully the next gen will provide more capable and flexible memory controllers, both so they can scale the top end for a full Pro-scale offering, and so there is more memory flexibility at the lower end e.g. the ability to get an M1 with 32+GB RAM, or a Pro with 64+.
We need to be able to run it properly in a Windows ARM VM on the M1 chips!
in traditional PC building, the CPU is quite distinct from the GPU. can anyone ELI5 what the benefits are to having the CPU closely integrated with GPU like the M1 has? seems a bit unwieldy but i dont know anything about computer architecture
I believe some Intel/AMD low power chips (non-performance laptops) have use this unified memory model as well.
But it became extremely common on phones where there was no historical baggage and it was thought out from day one. I believe all the consoles use a unified memory layout now, but I’m not 100% sure.
The usual limitation is you’re stuck with the on-die GPU, which can pale in comparison to a top of the line AMD or nVidia board.
https://www.theguardian.com/technology/2021/dec/07/apple-chi...
Data centers are also pretty conscious of power consumption, more power means more cooling infra required and higher energy bill, while it is not the top priority it certainly is a significant factor in decision making.
Good luck with those against Linux/FreeBSD workloads.
There are plenty of applications M1 is really good at, if Apple wanted to get into this space they could target , for example render farms.
Check your replies in 10 years and I'll be able to list a dozen ;P
But sarcasm aside yeah this chip looks insane.
For graphic card I don't try to argue because fps on Mac are very inferior in games than a average modern card. It's not even on the same league.
The M1 is as fast as the 1650. I'm getting great frame rates at 1440P High on X-Plane
> It took less than 6month to have a faster amd / intel CPU than the m1 back then, Apple charts are showing performance / watt which for a desktop PC is kind of irrelevant. In pure speed amd / intel are faster or will be very soon.
Perf/Watt is very relevant. Electricity costs money and you also want a cool room.
Of course Apple chips won’t work well for gaming, but what other benchmarks will this $2000 desktop win?
I'm almost certain that's not true, especially for machines that would compete with the Studio
> on PC you can get 2x the speed for 2x less the price.
Citation needed. This hasn't been true for a long time as far as I can tell.
Your $400 CPU needs at least another $1000 in parts just to boot (and those aren't even the parts you likely want to pair with it).
Your cost comparison is silly. Nobody compares singular CPUs to entire machines.
M1 is clearly the best design on the market for mobile devices and is merely very good for desktops. Let's keep the enthusiasm realistic.
Unsurprisingly, it's called Mac Studio, as in music studio, or art studio, or what have you studio, where these things matter.
This is a machine aimed at content creators.
I care. I work from home and my main power sink is my desktop. Considering the soaring energy prices these days I really do care about what my usage is.
Naming scheme aside, this is great!
https://9to5mac.com/2021/11/10/m1-pro-macbook-pro-cryptocurr...
M1 Pro -> 5.8 MH/s, with a 17w draw, means $12.82 a month profit. I don't imagine the M1 Ultra is too much better, maybe 20 MH/s at absolute most, but we'll see. It definitely won't be as economical as 3070 or 3080 FE cards at current profitability levels.
also note that mining calculator they used assumes 2 Ether per block paid to miners
In Ethereum it can be much much higher because people pay to use that blockchain. Mining can be insanely profitable and I’m not aware of any calculator that shows it. Everyone is operating on bad data. A cursory look right now shows latest blocks having 2.52 Ether in them, which is 26% greater yield.
Block 14348267 a few minutes ago had 4.83 Ether, 140% greater yield
There have been prolonged periods of time, weeks and months, where block rewards were 6-9 Ether.
Miners were raking it all in while the calculators said “2 Ether”
All this to say it could probably make $20-30 a month.
From the Mac Studio technical specifications
> Simultaneously supports up to five displays:
> Support for up to four Pro Display XDRs (6K resolution at 60Hz and over a billion colors) over USB-C and one 4K display (4K resolution at 60Hz and over a billion colors) over HDMI
Nice! Good enough to run a Solana node!
I was slightly annoyed that the M1 Max’s 64gb RAM puts it just under the system requirements, at that premium price
But I don't have any other theoretical use case for that much resources
Back in the Pentium-4 days, iirc I was able to get almost $250k in grants and $1.5M in subsidized loans to do accelerated refresh of a PC fleet and small datacenter, all through a utility's peak load reduction program.
Acceleration of that cycle with a Mac replacement, which usually has a 50-60 month lifecycle, is a pretty significant savings.