Heh, but it can't be that, no reason to think llms can count brackets needing a close any more than they can count words.
Consider the fact that GPT-4 can generate valid XML (meaning balanced tags, quotes etc) in base64-encoded form. Without CoT, just direct output.
I don't know what model Copilot uses these days, but it constantly makes bracket mistakes in Python.