Contrasting Intel AMX and Apple AMX
corsix.org
corsix.org
Sapphire rapids might well end up only shipping in Aurora and then being replaced immediately by it's successor.
To keep the record straight, we have: 1. AMX, 2. AVX, and 3. VMX.
What does that mean? Someone pointed out below Apple ships accelerate framework as a higher level supported mechanism to use these instructions?
Intel/AMD have been always good at documenting most of their stuff so perhaps we will see proper supported ones whenever Intel ships it.
What can you do easily and what’s hard?
I also expected to read something about relative performance of the two.
Which (assuming 1 multiply-add = 2 operations) is the same int8 operations/cycle as the Apple M1's float16 operations/cycle. Intel's 16-bit operations might be the same rate, but I'd guess half? That'll almost certainly be at a higher clock-speed, and one-per-core rather than one-per-four-P-cores. (And I think Apple might have doubled their throughput in M2. As you said, performance comparison is hard.)