I would speculate it’s struggling because of the linear nature of its output, and the red-herring words which crossover between categories.
Because the model can’t “look ahead”, it starts spitting out valid combinations, but without being able to anticipate that committing to a certain combination early on will lead to a mistake later.
I expect if you asked it to correct its output in a followup message, it could do so without much difficulty.