Jensen Huang must be having the time of his life right now. Nvidia's relationship with Apple went from pariah to prodigal son, real fast.
Jensen Huang must be having the time of his life right now. Nvidia's relationship with Apple went from pariah to prodigal son, real fast.
Edit: unless I misunderstood and they meant only inference.
people on HN routinely seem to overestimate Apple's capabilities
e: in fact, iirc just last month Apple released a paper unveiling their 'OpenElm' language models and they were all trained on nvidia hardware
I'm sure OpenAI will need to beef up their hardware to handle these requests - even as filtered down as they are - coming from all of the Apple users that will now be prompting calls to ChatGPT.
If Apple currently ships a single product with better AI performance-per-watt than Blackwell, I will eat my hat.
> These models run on servers powered by Apple silicon [...]
That doesn't mean that there are no Nvidia GPUs in these servers, of course.
Nvidia's GPUs are complex. They have a lot of dedicated, multipurpose acceleration hardware inside of them, and then they use CUDA to tie all those pieces together. Apple's GPUs are kinda the opposite way; they're extremely simple and optimized for low-power raster compute. Which isn't bad at all, for mobile! It just gimps them design-wise when they go up against purpose-built accelerators.
If we see Apple do custom Apple Silicon for the datacenter, it will be a pretty radically new design. The first thing they need is good networking; a full-size Nvidia cluster will use Mellanox Infiniband to connect dozens of servers at Tb/s speeds. So Apple would need a similar connectivity solution, at least to compete. The GPU would need to be bigger and probably higher-wattage, and the CPU should really emphasize core count over single-threaded performance. If they play their cards right there, they would have an Apple Silicon competitor to the Grace superchip and GB200 GPU.
No they don't. They say that the Secure Enclave participates in the secure boot chain, and in generating non-exportable keys used for secured transport. It reads to me as though user devices will encrypt requests to the keys held in the Secure Enclave of a subset of PCC nodes. A PCC node that receives the encrypted request will use the Secure Enclave to decrypt the payload. At that point, the general-purpose Application Processor in the PCC node has a cleartext copy of the user request for doing the needful inference, which _could_ be done on an NVidia GPU, but appears to be done on general-purpose Apple Silicon.
There is no suggestion that the user request is processed entirely within the Secure Enclave. The Secure Enclave is a cryptographic coprocessor. It almost certainly doesn't have the grunt to do inference.
They are the first ones to ship on-device inference at scale on non-nvidia hardware. Apple also has the means to build data center training hardware using apple silicon if they want to do so.
If they are serious about the OAI partnership they could also start to supply them with cloud inference hardware and strongarm them into only using apple servers to serve iOS requests.
i'm seeing people all over this thread saying stuff like that, it reads like fantasyland to me. Apple doesn't have the talent or the chips or suppliers or really any of the capabilities to do this, where are people getting it from?
Their ability to design a chip and networking fabric which is fast/efficient at training a narrow set of model architecture is not far fetched by any means.
> If they are serious about the OAI partnership they could also start to supply them with cloud inference hardware and strongarm them into only using apple servers to serve iOS requests
Apple addressed both these points in today’s preso.
1. They will send requests that require larger contexts to their own Apple Silicon-based servers that will provide Apple devices a new product platform called Private Cloud Compute.
2. Apple’s OS generative AI request APIs won’t even talk to cloud compute resources that do not attest to infrastructure that has a publicly available privacy audit.
You’re absolutely right. I got too excited about Apple’s strategy to encourage developers to use Apple Private Cloud Compute.
The UX for ChatGPT as shown for iOS 18 makes it obvious that you are sending data outside the Apple Silicon walled garden.
Which is neat, but it's not CUDA. It's an application-specific accelerator good at a small subset of operations, controlled by a high-level library the industry is unfamiliar with and too underpowered to run LLMs or image generators. The NPU is a novelty, and today's presentation more-or-less confirmed how useless it is for rich local-only operations.
> Apple also has the means to build data center training hardware using apple silicon if they want to do so.
They could, but that's not a competitor against an NVL72 with hundreds of terabytes of unified GPU memory. And then they would need a CUDA competitor, which could either mean reviving OpenCL's rotting corpse, adopting Tensorflow/Pytorch like a sane and well-reasoned company, or reinventing the wheel with an extra library/Accelerate Framework/MPS solution that nobody knows about and has to convert models to use.
So they can make servers, but Xserve showed us pretty clearly that you can lead a sysadmin to MacOS but you can't make them use it.
> they could also start to supply them with cloud inference hardware and strongarm them into only using apple servers to serve iOS requests.
I wonder how much money they would lose doing that, over just using the industry-standard Nvidia servers. Once you factor in the margins they would have made selling those chips as consumer systems, it's probably in the tens-of-millions.
Users absolutely don't care if their prompt response has been generated by a CUDA kernel or some poorly documented apple specific silicon a poor team at cupertino almost lost their sanity to while porting the model.
And haven't they already spent quite a bit on money on their pytorch-like MLX framework?
They most certainly will. If you run GPT-4o on an iPhone with MLX, it will suck. Users will tell you it sucks, and they won't do so in developer-specific terms.
The entire point of this thread is that Apple can't make users happy with their Neural Engine. They require a stopgap cloud solution to make up for the lack of local power on iPhone.
> And haven't they already spent quite a bit on money on their pytorch-like MLX framework?
As well as Accelerate Framework, Metal Performance Shaders and previously, OpenCL. Apple can't decide where to focus their efforts, least of which in a way that threatens CUDA as a platform.
Inference can run without it, and could so for years via ONNX. Now we are starting to see more back-ends becoming available.
But the point stands, these systems occupy a niche that Apple Silicon is poorly suited to filling. They run normal Linux, they support common APIs, and network to dozens of other machines using Infiniband.
This is Apple's favorite thing in the world. They already have an Apple-Silicon-only ML framework as of a few months ago, called MLX. Does anyone know about it? No. Do you need to convert models to use it? Yes.
It's not especially big news for Nvidia at all.
I just don't understand how they can compete on their own merits without purpose-built silicon; the M2 Ultra doesn't shine a candle to a single GB200. Once you consider how Nvidia's offerings are networked with Mellanox and CUDA universal memory, it feels like the only advantage Apple has in the space is setting their own prices. If they want to be competitive, I don't think they're going to be training Apple models on Apple Silicon.
NASDAQ average P:E - 31
NVidia's P:E - 71
That's a market of 1 vendor. That's ripe for attack.
You see, I want to live in a world where GPU manufacturers aren't perpetually hostile against each other. Even Nvidia would, judging by their decorum with Khronos. Unfortunately, some manufacturers would rather watch the world burn than work together for the common good. Even if a perfect CUDA replacement existed like it did with DXVK and DirectX, Apple will ignore and deny it while marketing something else to their customers. We've watched this happen for years, and it's why MacOS perennially cannot run many games or reliably support Open Source software. It is because Apple is an unreasonably fickle OEM, and their users constantly pay the price for Apple's arbitrary and unnecessary isolationism.
Apple thinks they can disrupt AI? It's going to be like watching Stalin try to disrupt Wal-Mart.
That's entirely the fault of AMD and Intel fumbling the ball in front of the other team's goal.
For ages the only accelerated backend supported by PyTorch and TF was CUDA. Whose fault was that? Then there was buggy support for a subset of operations for a while. Then everyone stopped caring.
Why I think it will go different this time: nVidia's competitors seem to have finally woken up and realized they need to support high level ML frameworks. "Apple Silicon" is essentially fully supported by PyTorch these days (via the "mps" backend). I've heard OpenCL works well now too, but have no hardware to test it on.
it's just a monopoly [1] , how hard can it be?
/s
- [1] practically, because of how widespread cuda is
…though it took two solid decades to even make a dent in x86.
If you're saying Apple is going to 'take on nvidia' in edge inference, then I don't disagree but I would hardly even count that as taking on nvidia.
It took almost a decade but the PA Semi acquisition showed that Apple was able to get out of the shadow of its PowerPC era.
Nvidia will remain a leader in this space for a long time. But things are going to play out wonky and Apple, when determined, are actually pretty good at executing on longer-term roadmaps.
I think Apple is going to make rapid and substantial advancements in on-device AI-specific hardware. I also think nVIDIA is going to continue to dominate the cloud infrastructure space for training foundational models for the foreseeable future, and serving user-facing LLM workloads for a long time as well.
You seem to be positioning this as a Ford vs Chevy duel, when (to me at least) the comparison should be to Ford vs Exxon.
Nvidia is an infrastructure company. And a darned good one. Apple is a user facing company and has outsourced infrastructure for decades (AWS & Azure being two of the well known ones).
The only way it hurts Nvidia is if Apple becomes the runaway leader of the pc market. Even then, Apple hasn’t shown any intent of selling GPUs or AI processors to the likes of AWS, or Azure or Oracle, etc.
Nvidia has a much bigger threat from with Intel/AMD or the cloud providers backward integrating and then not buying Nvidia chips. Again, no signs that Apple wants to do this.
For that they would have to develop servers that has a mass amount of whatever it is or sell the chips in the same manner Nvidia does today.
I dont see that future for Apple.
Microsoft / Google / or other major cloud companies would do extremely well if they could develop it and just keep it as a major win for their cloud products.
Azure is running OpenAI as far as I have heard.
Imagine if M$ made a crazy fast GPU/whatever. It would be a huge competitive advantage.
Can it happen? I dont think so.
So all in all, not sure if it's that great for Nvidia.
What would not have been positive for Nvidia is Apple saying they've adapted their HW to server chips and would be partnering with OpenAI to leverage them, but that didn't happen. Apple is busy handing cash back to investors and not seriously pursuing anything but inference.
current situation is like nvidia devs using macs at work giving mr cook some satisfaction or something.