This is obviously intended for running local AI, but the memory bandwidth of 300 GB/sec seems a limiting factor, since this is what limits single user LLM inference tokens/sec. For comparison the M2 Ultra does up to 800 GB/sec, thanks to having a 1024-bit bus vs the RTX Spark's 256-bit bus.