Point 1 doesn't feel like a good enough reason. The number of tokens outputted as a JSON is so small if you tell GPT to output it properly.
I think what factors in even more when you use the API is that you do not have fine-grained control over the generation process. If you follow the MS guidance approach, you fill in structured text yourself, and then let the model generate only the value parts, e.g. up to the next quote. To do that more or less word by word, you have multiple API calls, and have to be very smart about providing the right stop tokens.