I need more coffee too early!
I need more coffee too early!
A hard restriction would be a regex or a simpler model checking your prompt for known or suspected bad prompts and refusing outright.
If NLP was that easy, we wouldn't have needed to invent transformer models, and we'd have had things as capable as ChatGPT about the same time that Microsoft was selling Encarta on CD.
The reality is, this soft fuzzy thing is the only practical way to minimise the Scunthorpe problem (and its equivalents for false negatives): https://en.wikipedia.org/wiki/Scunthorpe_problem
Could be function calling, could be some other mechanism. It seems rather trivial to restrict the system at that interface point where the NLP gets translated into image generation instructions and simply cut it off at a limit.
Since there is an interface point where the NLP is translated into some image generation call path, presumably, the response generation can see that it only got 1 image back instead of n. Even a system prompt could be added to normalize the output text portion of the response.
Alternatively, if ChatGPT does think it is suspicous that it only got one URL, it might end up responding with "the system seems to be not working right now, I'm only getting partial results for your query", because it doesn't know that the system is only going to return a single image.
This is getting into speculative territory, so I guess the true answer could also be "OpenAI are amateurs are prompting ChatGPT", but it seems less likely.