Apple M3 Ultra
apple.com
apple.com
The the question is if a llm will run with usable performance at that scale? The point is there's diminishing returns despite having enough uRAM with the same amount of memory bandwidth even with increased processing speed of the new chip for AI.
So there must be a min-max performance ratio between memory bandwidth and the size of the memory pool in relation to the processing power.
From my napkin math, the M3 Ultra TFLOPs is still relatively low (around 43 FP16 TFLOPs?), but it should be more than enough to handle bs=1 token generation (should be way <10 FLOPs/byte for inference). Now as far is its prefill/prompt processing speed... well, that's another matter.
This is the big question to have answered. Many people claim Apple can now reliably be used as a ML workstation, but from the numbers I've seen from benchmarks, the models may fit in memory, but the performance for tok/sec is so slow to not feel worth it, compared to running it on NVIDIA hardware.
Although it be expensive as hell to get 512GB of VRAM with NVIDIA today, maybe moves like this from Apple could push down the prices at least a little bit.
For the self-attention mechanism, memory bandwidth requirements scale ~quadratically with the sequence length.
If they didn't increase the memory bandwidth, then 512GB will enable longer context lengths and that's about it right? No speedups
For any speedups You may need some new variant of FlashAttention3 or something along similar lines to be purpose built for Apple GPUs.
(and the 512GB version is $4,000 more rather than $10,000 - that's still worth mocking, but it's nowhere near as much)
On dual (SP5) Epyc I believe the memory bandwidth is somewhat greater than this apple product too... and at apple's price points you can have about twice the ram too.
Presumably the apple solution is more power efficient.
As for practicality, which mainstream applications would benefit from this much memory paired with a nice but relative mid compute? At this price-point (14K for a full specced system), would you prefer it over e.g. a couple of NVIDIA project DIGITS (assuming that arrives on time and for around the announced the 3K price-point)?
I would have assumed you’d want to save the best process/node for processing, and could use a less expensive processes for RAM.
The pricing isn't as insane as you'd think, 96 to 256GB is 1500 which isn't 'cheap' but, it could be worse.
All in 5,500 gets you a ultra with 256GB memory, 28 cores, 60 GPU cores, 10Gb network - I think you'd be hard pushed to build a server for less.
And the storage solution still makes no sense of course, a machine like this should start at 4TB for $0 extra, 8TB for $500 more, and 16TB for $1000 more. Not start at a useless 1TB, with the 8TB version costing an extra $2400 and 16TB a truly idiotic $4600. If Sabrent can make and sell 8TB m.2 NVMe drives for $1000, SoC storage should set you back half that, not over double that.
price premium probably, but chip lithography errors (thus, yields) at the huge memory density might be partially driving up the cost for huge memory.
funny that people think this is so new, when CRAY had Global Heap eons ago...
If you are going to argue that the OS or even below that the hardware could be compromised to still enable exfiltration, that is true, but it is a whole different ballgame from using an external SaaS no matter what the service guarantees.
A long time ago Apple had a rackmount server called Xserve, but there’s no sign that they’re interested in updating that for the AI age.
That Said, 512GB of unified ram with access to the NPU is absolutely a game changer. My guess is that Apple developed this chip for their internal AI efforts, and are now at the point where they are releasing it publicly for others to use. They really need a 2U rack form for this though.
This hardware is really being held back by the operating system at this point.
The CPUs have zero competition in terms of speed, memory bandwidth. Still blown away no other company has been able to produce Arm server chips that can compete.
I’d be curious to know if this changes that. It’d take a lot more than doubling cores to take out the very high power AMD parts, but this might squeeze them a bit.
Interestingly, AMD has also been investing heavily in unified RAM. I wonder if they have / plan an SoC that competes 1:1 with this. (Most of the parts I’m referring to are set up for discrete graphics.)
It reminds me of the 1990s when my old school was using Sun machines based on the 68k series and later SPARC and we were blown away with the toaster-sized HP PA RISC machine that was used for student work for all the CS classes.
Then Linux came out and it was clear the 386 trashed them all in terms of value and as we got the 486 and 586 and further generations, the Intel architecture trashed them in every respect.
The story then was that Intel was making more parts than anybody else so nobody else could afford to keep up the investment.
The same is happening with parts for phones and TSMC's manufacturing dominance -- and today with chiplets you can build up things like the M3 Ultra out of smaller parts.
Maybe not at the same power consumption, but I'm sure mid-range Xeons and EPYCs mop the floor with the M3 Ultra in CPU performance. What the M3 Ultra has that nobody else comes close is a decent GPU near a pool of half a terabyte of RAM.
FYI Apple runs Linux in their DC, so no Apple hardware in their own servers.
Our business "only" sees about 1,000-25,000 req/min, our message brokers transmit MAX 25k msg/s. Easily handled by a rack of 10 servers for redundancy.
We are not Google and we don't pretend to be, so we don't care about power, as the difference is a few dollars a month.
It really is. Even if they themselves won't bring back their old XServe OS variant, I'd really appreciate it if they at least partnered with a Linux or BSD (good callout, ryao) dev to bring a server OS to the hardware stack. The consumer OS, while still better (to my subjective tastes) than Windows, is increasingly hampered by bloat and cruft that make it untenable for production server workloads, at least to my subjective standards.
A server OS that just treats the underlying hardware like a hypervisor would, making the various components attachable or shareable to VMs and Containers on top, would make these things incredibly valuable in smaller datacenters or Edge use cases. Having an on-prem NPU with that much RAM would be a godsend for local AI acceleration among a shared userbase on the LAN.
With all my love and respect for "Apple rumors" writers; this was always "I read five blogposts about CPU design and now I'm an expert!" territory.
The speculation was based on the M3 Maxes die shots not having the interposer visible, which... implies basically nothing whether that _could have_ been supported in an M3 Ultra configuration; as evidenced by the announcement today.
No M3 has thunderbolt 5.
This is a new chip with M3 marketing. I’d expect this from Intel, not Apple.
Apple could either create a 2U rack hardware and support Linux (and I mean Apple supporting it, not hobbysts), or have a build of Darwin headless that could run on that hardware. But in the later case, we probably wouldn't have much software available (though I am sure people would eventually starting porting software to it, there is already MacPorts and Homebrew and I am sure they could be adapted to eventually run in that platform).
But Apple is also not interested in that market, so this will probably never happen.
they're just a tiny company with shareholders who are really tired of never earning back their investments. give 'em a break. I mean they're still so small that they must protect themselves by requiring that macs be used for publishing iPhone and iPad applications.
M1 Max - 24 to 32 GPU cores
M2 Max - 30 to 38 GPU cores
M3 Max - 30 to 40 GPU cores
M4 Max - 32 to 40 GPU cores
I also looked up the announcement dates for the Max and the Ultra variant in each generation.
M1 Max - October 18, 2021
M1 Ultra - March 8, 2022
M2 Max - January 17, 2023
M2 Ultra - June 5, 2023
M3 Max - October 30, 2023
M3 Ultra - March 12, 2025
M4 Max - October 30, 2024
> My guess is that Apple developed this chip for their internal AI efforts
As good a guess as any, given the additional delay between the M3 Max and Ultra being made available to the public.
For SMBs or Edge deployments where redundancy isn't as critical or budgets aren't as large, this is an incredibly compelling offering...if Apple actually had a competent server OS to layer on top of that hardware, which it does not.
If they did, though...whew, I'd be quaking in my boots if I were the usual Enterprise hardware vendors. That's a damn frightening piece of competition.
It's worth adding the M3 Ultra has 819GB/s memory bandwidth [1]. For comparison the RTX 5090 is 1800GB/s [2]. That's still less but the M4 Mac Minis have 120-300GB/s and this will limit token throughput so 819GB/s is a vast improvement.
For $9500 you can buy a M3 Ultra Mac Studio with 512GB of unified memory. I think that has massive potential.
[1]: https://www.apple.com/mac-studio/specs/
[2]: https://www.nvidia.com/en-us/geforce/graphics-cards/50-serie...
https://digitalspaceport.com/how-to-run-deepseek-r1-671b-ful...
between 4.25 to 3.5 TPS (tokens per second) on the Q4 671b full model.
3.5 - 4.25 tokens/s. You're torturing yourself. Especially with a reasoning model.This will run it at 40 tokens/s based on rough calculation. Q4 quant. 37b active parameters.
5x higher price for 10x higher performance.
I just want a break from MacOS, I'll be buying a Thinkpad and will probably never come back. This isn't my moaning, I understand it's their market, but if their hardware supported Linux (especially dual booting) or Docker native, I'd probably be buying Apple for the next decade and now I just won't be.
I love Apple but they love to speak in half truths in product launches. Are they saying the M3 Ultra is their first Thunderbolt 5 computer? I don't recall seeing any previous announcements.
I assume that there's a community of developers focusing on leveraging this hardware instead of complaining about the operating system.
> Apple’s custom-built UltraFusion packaging technology uses an embedded silicon interposer that connects two M3 Max dies across more than 10,000 signals, providing over 2.5TB/s of low-latency interprocessor bandwidth, and making M3 Ultra appear as a single chip to software.
The comment was that the press had reported that the interposer wasn't available. This obviously uses some form of interposer, so the question is if the press missed it, or Apple has something new.
Please elucidate.
^ has a lot of elaborations on this subject
The Apple ecosystem is a walled garden.
what internal AI efforts?
Apple Intelligence is bunkers, and Apple MLX framework remains a hobby project for Apple
It’s their spin of the Google strategy of targeting providjng services to their enterprise GCP customer. I think we’ll see more out of them long term.
They now bump it to 512GB. Along with insane price tag of $9499 for 512GB Mac Studio. I am pretty sure this is some AI Gold rush.
A 4-bit quantization of Llama-3.1 405b, for example, should fit nicely.
When running LLMs on Docker with an Apple M3 or M4 chip, they will operate in CPU mode regardless of the chip's class, as Docker only supports Nvidia and Radeon GPUs.
If you're developing LLMs on Docker, consider getting a Framework laptop with an Nvidia or Radeon GPU instead.
Source: I develop an AI agent framework that runs LLMs inside Docker on an M3 Max (https://kdeps.com).
Additionally, I would assume this is a very low-volume product, so it being on N3B isn't a dealbreaker. At the same time, these chips must be very expensive to make, so tying them with luxury-priced RAM makes some kind of sense.
Makes it even more puzzling what they are doing with the M2 Mac Pro.
[0] https://www.numerama.com/tech/1919213-m4-max-et-m3-ultra-let...
[1] More context on Macrumors: https://www.macrumors.com/2025/03/05/apple-confirms-m4-max-l...
And anyway, I think the M2 Mac Pro was Apple asking customers "hey, can you do anything interesting with these PCIe slots? because we can't think of anything outside of connectivity expansion really"
RIP Mac Pro unless they redesign Apple Silicon to allow for upgradeable GPUs.
Either that or kill the Mac Pro altogether, the current iteration is such a half-assed design and blatantly terrible value compared to the Studio that it feels like an end-of-the-road product just meant to tide PCIe users over until they can migrate everything to Thunderbolt.
They recycled a design meant to accommodate multiple beefy GPUs even though GPUs are no longer supported, so most of the cooling and power delivery is vestigial. Plus the PCIe expansion was quietly downgraded, Apple Silicon doesn't have a ton of PCIe lanes so the slots are heavily oversubscribed with PCIe switches.
(sorry, should have specified that the NPU and GPU cores need to access that ram and have reasonable performance). I specified it above, but people didn't read that :-)
Compared to Nvidia's Project DIGITS which is supposed to cost $3K and be available "soon", you can get a specs matching 128GB & 4TB version of this Mac for about $4700 and the difference would be that you can actually get it in a week and will run macOS(no idea how much performance difference to expect).
I can't wait to see someone testing the full DeepSeek model on this, maybe this would be the first little companion AI device that you can fully own and can do whatever you like with it, hassle-free.
at 819 GB per second bandwidth, the experience would be terrible
[1] Asus just announced the world’s first Thunderbolt 5 eGPU:
https://www.theverge.com/24336135/asus-thunderbolt-5-externa...
When I connect to my Mac Studio via Macbook I can select that mode, then change the Displays setting to Dynamic Resolution and then my 'thin client':
- Is fullscreen using the entire 16:10 Macbook screen
- Gets 60 fps low latency performance (including on actual games)
- Transfers audio, I can attend meetings in this mode
- Blanks the host Mac Studio screen
All things that were impossible via VNC - RDP is much better but this new High Performance Screen Share is even more powerful.
The thin lightweight laptop that remotes into a loaded machine has always been my idea of high mobility instead of suffering a laptop running everything locally. This works via LTE as well with some firewall setup.
The hardware has evolved faster than software at Apple. It’s usually the opposite with most tech companies where hardware is unable to keep up with software.
That said, there are efforts being made to use the NPU. See: https://github.com/Anemll/Anemll - you can now run small models directly on your Apple Silicon Mac's NPU.
It doesn't give better performance but it's massively more power efficient than using the GPU.
The Neural Engine is useful for a bunch of Apple features, but seems weirdly useless for any LLM stuff... been wondering if they'd address it on any of these upcoming products. AI is so hype right now it seems odd that they have specialised processor that doesn't get used for the kind of AI people are doing. I can see in the latest release:
> Mac Studio is a powerhouse for AI, capable of running large language models (LLMs) with over 600 billion parameters entirely in memory, thanks to its advanced GPU
https://www.apple.com/newsroom/2025/03/apple-unveils-new-mac...
i.e. LLMs still run on the GPU not the NPU
Not sure how much storage to get. I was floating the idea of getting less storage, and hooking it up to a TB5 NAS array of 2.5” SSDs, 10-20tb for models + datasets + my media library would be nice. Any recommendations for the best enclosure for that?
I also want to build the thing you want. There are no multi SSD M2 TB5 bays. I made one that holds 4 drives (16TB) at TB3 and even there the underlying drives are far faster than the cable.
My stuff is in OWC Express 4M2.
This is my understanding (probably incorrect in some places)
1. NVIDIA's big advantage is that they design the hardware (chips) and software (CUDA). But Apple also designs the hardware (chips) and software (Metal and MacOS).
2. CUDA has native support by AI libraries like PyTorch and Tensorflow, so works extra well during training and inference. It seems Metal is well supported by PyTorch, but not well supported by Tensorflow.
3. NVIDIA uses Linux rather than MacOS, making it easier in general to rack servers.
In terms of hardware - Apple designs their GPUs for GPU workloads, whereas Nvidia has a decades-old lead on optimizing for general-purpose compute. They've gotten really good at pipelining and keeping their raster performance competitive while also accelerating AI and ML. Meanwhile, Apple is directing most of their performance to just the raster stuff. They could pivot to an Nvidia-style design, but that would be pretty unprecedented (even if a seemingly correct decision).
And then there's CUDA. It's not really appropriate to compare it to Metal, both in feature scope and ease of use. CUDA has expansive support for AI/ML primatives and deeply integrated tensor/SM compute. Metal does boast some compute features, but you're expected to write most of the support yourself in the form of compute shaders. This is a pretty radical departure from the pre-rolled, almost "cargo cult" CUDA mentality.
The Linux shtick matters a tiny bit, but it's mostly a matter of convenience. If Apple hardware started getting competitive, there would be people considering the hardware regardless of the OS it runs.
I bought a refubished M3 max to run LLMs (can only go up to 70b with 4 bit quant), and it is only slightly slower than the more expensive M4 max.
I feel like I should be able to spend all my money to both get the fastest single core performance AND all the cores and available memory, but Apple has decided that we need to downgrade to "go wide". Annoying.
I'm a major Apple skeptic myself, but hasn't there always been a tradeoff between "fastest single core" vs "lots of cores" (and thus best multicore)?
For instance, I remember when you could buy an iMac with an i9 or whatever, with a higher clock speed and faster single core, or you could buy an iMac Pro with a Xeon with more cores, but the iMac (non-Pro) would beat it in a single core benchmark. Note: Though I used Macs as the example due to the simple product lines, I thought this was pretty much universal among all modern computers.
For M3 and M4 machines, hardware support is pretty derilict: https://asahilinux.org/docs/M3-Series-Feature-Support/
I'm not sure if this is me not maintaining it properly (e.g fans having dust block them) - but I've always got this sense that Apple throttles their older devices in some indirect ways. I experience it the most with iPhones - my old iPhone is pretty slow doing basic things despite nothing really changing on it (just the OS updating?)
So my only concern with this is - how many years until it's slow enough to annoy you into buying a new one?
"Just the OS updating" is not insignificant. Software developers, in general, are not known for making sure latest versions of their software run smoothly on older hardware.
Also, performance on iPhones is throttled when your battery is very old. There was a whole class-action lawsuit about it.
Do you expect this will be able to handle AI workloads well?
All I’ve heard for the past two years is how important a beefy GPU is. Curious if that holds true here too.
The model weights (billions of parameters) must be loaded into memory before you can use them.
Think of it like this: Even with a very fast chef (powerful CPU/GPU), if your kitchen counter (VRAM) is too small to lay out all the ingredients, cooking becomes inefficient or impossible.
Processing power still matters for speed once everything fits in memory, but it's secondary to having enough VRAM in the first place.
The thing with these Apple chips is that they have unified memory, where CPU and GPU use the same memory chips, which means that you can load huge models into RAM (no longer VRAM, because that doesn't exist on those devices). And while Apple's integrated GPU isn't as powerful as an Nvidia GPU, it is powerful enough for non-professional workloads and has the huge benefit of access to lots of memory.
These are unified memory. The M3 Ultra with 512gb has as much VRAM as sixteen 5090.
1. What are various average joe (as opposed to researchers, etc.) use cases for running powerful AI models locally vs. just using cloud AI. Privacy of course is a benefit, but it by itself may not justify upgrades for an average user. Or are we expecting that new innovation will lead to much more proliferation of AI and use cases that will make running locally more feasible?
2. With the amount of memory used jumping up, would there be a significant growth for companies making memories? If so, which ones would be the best positioned?
Thanks.
A local model will do anything you ask it to, as far as it "knows" about it. It doesn't need to please investors or be afraid of bad press.
LM Studio + a group of select models from huggingface and you can do whatever you want.
For generic coding assistance and knowledge, online services are still better quality.
Apple seems to be using LPDDR, but HBM will also likely be a key tech. SK Hynix and Samsung are the most reputable for both.
I think a great use case for this would be in a company that doesn't want all of their employees sending LLM queries about what they're working on outside the company. Buy one or two of these and give everybody a client to connect to it and hey presto you've got a secure private LLM everybody in the company can use while keeping data private.
I'm in Hong Kong, I can't even subscribe to OpenAI or Claude directly, though granted this doesn't so much apply to the already "open" models
I do not have a good sense of how well quality scales with narrow MoEs but even if we get something like Llama 3.3 70b in quality at only 8b active parameters people could do a ton locally.
If it is equivalent, then the machine pays for itself in 300 hours. That's incredible value.
I wonder if the plan is to only release Ultras for odd number generations.
Gamers don't generally use a mac because of the lack of games and I'm guessing those who are really into LLMs use Linux for the flexibility. Video editing can be done on much cheaper hardware.
Very rich LLM enthusiasts who wants to try out mac?
You can get a good experience on a Windows or Linux machine with DaVinci Resolve, but that’s mostly because of the way better GPUs like the 4090/RTX series you’ve got at your disposal.
Hah, I see what they did there.
It feels like one should be able to build a good machine for 3/4k if not less with 6 16GB mid level gaming GPUs.
Well, duh, it would be a shame if you made a step backwards, wouldn't it? I hate that stupid phrase...
This is a cool computer, but not something I'd want to lug around.
Hot take: You can tie yourself into six knots trying to spin a yarn about why the M3 Ultra spec is super awesome for some AI use-case, meanwhile you could buy a Mac Mini and like 200 million GPT-4o tokens for the cost of this machine that can't even run R1.
$9499
What ever happening to competition in computing?
Computing hardware competition used to be cut throat, drop dead, knife fight, last man standing brutally competitive. Now it's just a massive gold rush cash grab.
Could cost half of that and it would still be uninteresting for my use cases.
For AI, on-demand cloud processing is magnitudes better in speed and software compatibility anyway.
I'm curious what instruction sets may have been included with the M3 chip that the other two lack for AI.
So far the candidates seem to be NVIDIA digits, Framework Desktop, M1 64gb M2/M3 128gb studio/ultra.
The GPU market isn't competitive enough for the amount of VRAM needed. I was hoping for an Battlemage GPU Model with 24GB that would be reasonably priced and available.
The framework desktop and devices I think a second generation will be significantly better than what's currently on offer today. Rationale below...
For a max spec processor with ram at $2,000, this seems like a decent deal given today's market. However, this might age very fast for three reasons.
Reason 1: LPDDR6 may debut in the next year or two this could bring massive improvements to memory bandwidth and capacity for soldered on memory.
LPDDR6 vs LPDDR5 - Data bus width - 24 bits, 16 bits Burst length - 24 bits, 15 bits Memory bandwidth - Up to 38.4 GB/s, Up to 6.7 GB/s
- Camm ram may or may not be maintain signal integrity as memory bandwidth increases. Until I see it implemented for a AI use-case in a cost-effective manner, I am skeptical.
Reason 2: - It's a laptop chip with limited PCI lanes and reduced power envelope. Theoretically, a desktop chip could have better performance, more lanes, socketable (Although, I don't think I've seen a socketed CPU with soldered RAM)
Reason 3: In addition, what does hardware look like being repurposed in the future compared to alternatives?
- Unlike desktop or server counterparts which can have a higher cpu core count, PCEe/IO Expansion, this processor with its motherboard is limited on re-purposing later down the line as a server to self-host other software besides AI. I suppose could be turned into a overkill, NAS with ZFS and HBA Single Controller Card in new case.
- Buying into the framework desktop is pretty limited based on the form factor. Next generation might be able to include a 16x slot fully populated, a 10G nic. That seems about it if they're going to maintain the backward compatibility philosophy given the case form factor.
Did they say why there’s not an m4 ultra?
Soldered?
Figure out a way to make it unified without also soldering it, and you'll be a billionaire.
Or are you just grinding a tired, 20-year-old axe.
The issue is availability of chips and most likely you have to know which components to change so the new memory is recognised. For instance that could be changing a resistor to different value or bridging certain pads.
Is anyone other than a vanishingly small number of hard core hobbiests going to upgrade from an M4 to an M4 Ultra?
Has anyone has a ballpark number how many tokens per second we can get with this?
I thought it was few weeks ago when M4 Max came by.
Why? They have too many M3 chips on stock?
what's the point of 512GB RAM for LLMs on this Mac Studio if the speed is painfully slow?
it's as if Apple doesn't want to compete with Nvidia... this is really disappointing in a Mac Studio. FYI: M2 Ultra already has 800GB/s bandwidth
NVIDIA RTX 4080: ~717 GB/s
AMD Radeon RX 7900 XTX: ~960 GB/s
AMD Radeon RX 7900 XT: ~800 GB/s
How's that slow exactly ?
You can have 10000000Gb/s and without enough VRAM it's useless.
what's the point of 512GB RAM for LLMs on this Mac Studio if the speed is painfully slow?
You can fit the entire Deepseek 671B q4 into this computer and get 41 tokens/s because it's an MoE model.Even for its intended AI audience, the ISA additions in M4 brought significant uplift.
Are they waiting to put M4 Ultra into the Mac Pro?
please take my money now
With an M3 Ultra going into the Mac Studio, Apple could differentiate from the Mac Pro, which could then get the M4 Ultra. Right now, the Mac Studio and Mac Pro oddly both have the M2 Ultra and same overall performance.
https://x.com/markgurman/status/1896972586069942738In terms of software, recent NVIDIA and AMD research has focused on fast evaluation of small ~4 layer MLPs using FP8 weights for things like denoising, upscaling, radiance caching, and texture and material BRDF compression/decompression.
NVIDIA has just put out some new graphics API extensions and samples/demos for loading a chunk of neural net weights and performing inference from within a shader.
For comparison, a single consumer card like the RTX 5090 is only 32 GB of memory, has 1792 GB/s memory and 3593 TOPS of compute.
The use cases will be limited. While you can't run a 600B model directly like Apple says(cause you need more memory for that), you can run a quantized version, but it will be very slow unless its a MoE architecture.
The compute level you’re talking about on the M3 Ultra is the neural engine. Not including the GPU.
I expect the GPU here will be behind a 5090 for compute but not by the unrelated numbers you’re quoting. After all, the 5090 alone is multiple times the wattage of this SoC.
Thats going to be the NPU specifically. Pretty much nothing on llm front seems to use NPUs at this stage (copilot snapdragon laptops aside) so not sure the low number is a problem
It's nice that these devices have loads of memory, but they don't have remotely the necessary level of compute to be competitive in the AI space. As a fun thing to run a local LLM as a hobbyist, sure, but this presents zero threat to nvidia.
Apple hardware is irrelevant in the AI space, outside of making YouTube "I ran a quantized LLM on my 128GB Mac Mini" type content for clicks, and this release doesn't change that.
Looks like a great desktop chip though.
It would be nice if nvidia could start giving their less expensive offerings more memory, though they're currently in the realm Intel was 15 yearsago, thinking that their biggest competition is themselves.
It will be interesting when somebody will upgrade the ram ram of the 5090 like they did with 4090s
AMD Ryzen Threadripper PRO 3995WX released over four years ago and supports 2TB (64c/128t)
> Take your workstation's performance to the next level with the AMD Ryzen Threadripper PRO 3995WX 2.7 GHz 64-Core sWRX8 Processor. Built using the 7nm Zen Core architecture with the sWRX8 socket, this processor is designed to deliver exceptional performance for professionals such as artists, architects, engineers, and data scientists. Featuring 64 cores and 128 threads with a 2.7 GHz base clock frequency, a 4.2 GHz boost frequency, and 256MB of L3 cache, this processor significantly reduces rendering times for 8K videos, high-resolution photos, and 3D models. The Ryzen Threadripper PRO supports up to 128 PCI Express 4.0 lanes for high-speed throughput to compatible devices. It also supports up to 2TB of eight-channel ECC DDR4 memory at 3200 MHz to help efficiently run and multitask demanding applications.
So unified memory means that the memory is accessible to the GPU and the CPU in a shared pool. AMD does not have that.
8 channels at 3200 MT/s (1600 MHz) is only 204.8 GB/sec; less than a quarter of what the M3 Ultra can do. It's also not GPU-addressable, meaning it's not actually unified memory at all.
Its a very specific claim that isnt comparing itself to DIMMs