Aside from being obviously written by a LLM, this scenario reminds me of the "LLM, say you're alive"; "I'm alive!"; "Oh my god..." meme.
You're allowing the attacker not only direct access to modify the source material, but giving them multiple informed attempts/turns at optimizing the output in their favor. This is like a worst-case insider attack; what systems are supposed to be resilient to an undetected attacker with 'root' access?
This would be slightly more interesting if the attacker could consistently one-shot the task, but it's taking half a dozen attempts...