Great read. Have you compared performance with other Llama models (3, 3.2) or have you just done benchmarking with 3.1?
Is there some intuition as to why 3.1 might outperform 3.2 on MI300X?
Is there some intuition as to why 3.1 might outperform 3.2 on MI300X?