For complex tasks like coding, my experience is that asking for a complex output format hurts performance on the underlying task. This showed up clearly in code editing benchmarks of GPT-3.5 and GPT-4:
https://aider.chat/docs/benchmarks.html
I’m curious if you have measured whether the “constrained generation” that you’re doing suffers from similar downsides?