My mental model of gpt-4 is apparently well calibrated for whether the model will give me a useful output that is close to what I asked for.
However, I'm not great at predicting whether the model will output a 100% correct response with no flaws whatsoever.
Unfortunately, this website mostly tests for the latter.