Oh, I see. That must be frustrating to folks at OpenAI. Their product rests on the quality of their models, and making users unable to see which results came from their best doesn't help.
FWIW, GPT-4 and GPT-4 Turbo via developer API call both seem to produce the result you expect.