Theoretically awesome, but this might have some interesting market consequences for everyone else.
Theoretically awesome, but this might have some interesting market consequences for everyone else.
256GB/sec, so roughly M4 Pro throughput.
Or a dual socket AMD Turin.
Or a grace+hopper (assuming you offload to the hopper).
Latency is dictated by the laws of physics, more bandwidth is easy, but not cheap.
You could run R1 671B using unsloth’s quantized version that fits in <80GB. Not sure why that would be a benchmark though, there’s nothing that can run the model at full precision right now except for (very slow) server hardware.
Not if you are using the CPU. I am under the impression most inference use cases are memory bandwidth limited, not compute limited, so running on the GPU would gain you little to nothing unless the GPU has faster access to the shared memory.