Intel Is Counting on AI Inference to Save the Xeon CPU
nextplatform.com
nextplatform.com
I tried Mistrel 7B in MlStudio.AI on an old HP z440 with 4x16=64GB DDR4 Reg RAM. It bottlenecks on the RAM, with the E5-1630 v4 running at 50%. It prints text faster than I can read.
If I were to design a processor for LLM, I would throw in a lot of cores, a lot of cache, and a lot of High Bandwidth Memory on the chip. This is basically the newer Xeons, so I think they are headed in the right direction.
I wish Intel had put a lot more HBM on their Xeon Phi accelerator cards. If they had it would have made them perfect for LLMs.