I was impressed by their _CPU_ kernels in OpenVINO which I studied very closely after discovering I cannot beat their performance even on kernels specialized for my specific tensor shapes. Turns out the reason I can't beat them is because they jitted their kernels with xbyak, and specialized them for those shapes, too. They also obviously knew exactly how to maximize the ALU utilization, and counted every cycle, something that's very hard to do unless you specialize _just_ on CPU performance and have access to the internal documentation. Top notch work. But that team used to be in Russia back then, in Nizhny Novgorod, IDK if it's still a part of Intel.