One related experiment, however, already suggests the result. Someone tried to get ChatGPT to solve advent of code challenges: https://github.com/golergka/advent-of-code-2022-with-chat-gp.... These challenges are very clear, and it already seems to struggle to get the answers. One the second day, it took 12 tries to get it right. If it struggles this much with clear requirements, I don't think it'll do well with vague ones.