It's like accusing somebody of being a lousy chef because they have such terrible taste in takeout. There's just no connection between these things.
Look, there is a lot of pro-AI nonsense going around right now that I too would like to shut down, but none of that changes the fact that to prove that a machine is incapable of a task requires an experimental structure in which the machine is performing as best as it possibly can and still comes up short, or an argument based on fundamentals about what the machine is (e.g. a steam engine can't move faster than the speed of sound in steam).
No pile of failures, however embarrassing, can prove the impossibility of success, that's just not how evidence works.
Or if you use AI to start removing the autographs in the training data before givig it another go then that is not because it got there trough reasoning. If you adjust the harness or somehow add rules to the prompt to keep it within the limits of what we know people probably want then the improved result is not because it got there trough reasoning.
Unless you take a fundamentally different approach then underlying ways that caused it to add the autograph or add the plastic or what have you remain same and we know that. There is evidence for that.