Like, the post is about how you can do GEMM tuning on AMD GPUs, a subject which is inherently super interesting–there's a lot of nuance to writing optimal kernels, and some of this is expressed in the article, too. Combine that with an architecture that isn't Nvidia? It's an excellent setup for something that would be interesting to read.
Which makes the actual conclusion all the more disappointing, IMO. There's nothing about what the actual optimizations are. It's just "oh yeah we tuned our GEMMs and now LLaMA is faster". Like, I get that nobody actually cares about GEMM and they just want tokens to come out of their GPU. But still, that's like writing a blog post about how you can speed up your game with SIMD and then posting some charts of how Cyberpunk 2077 gives you 2x the frame rate now. Ok, but how? I just feel like the interesting part is missing.