Very interesting and Would love to see the experiments. Quick question: what do you mean about kernel dependent ?
We saw different results of pipelining with the Attention kernel vs the MLP kernel (since MLP W1 has to project the attention results into a much higher dimension, the arithmetic intensity shifts towards compute bound characteristics)