Eg you can relatively easy hack up a bit of code to create questions at random. At the most primitive, you just have a simple template that you fill in randomly. Like 'If I put _a down in front of _b but behind _c, what item will be in the middle?' with various _a, _b and _c.
If you make it slightly more complicated and have big enough pools to draw from, you can guarantee that the questions you are generating were not in the training set: even if just because you can sample from, say, 10^100 different questions pretty easily, and I'm fairly sure their training set was smaller than that.
The key question is where the boundaries are. Maybe they should be part of the response - a per sentence or per paragraph "confidence scale" that signals how hard they extrapolated from their trained space (I know transformers work per token, but sentence/paragraph would be better human UX).
Of course, if they were trained on garbage input, that would only tell you how accurately they sticked to the garbage. But it would still be invaluable instrumentation for the end user, not to mention for the API provider. They could look at high demand subjects with low confidence answers and prioritize that for further training.