You're fundamentally describing a stochastic parrot here, and not an intelligence. When you say that humans can also make the same mistake, you're ignoring the fact that every human testing these systems and finding it lacking are also comparing their experience to a lifetime of interacting with human intelligence. Every anecdote of this sort is an example, typically based on repeatedly trying an assortment of prompts (eliminating two of your variables -- random values & varying prompts), on a large variety of tasks, which is curated down to a single example for the sake of brevity.
To say that "oh maybe it would have gotten the right answer if you got lucky or tried harder or twiddled a knob that you don't have access to" is simply nowhere near the extraordinary evidence that is required to prove the extraordinary claim of "real understanding." The reason that you don't know what to do about these anecdotes is that you lack the evidence to properly rebut them. You can only wave your hands at ill-defined properties and gaslight about the user holding it wrong. Perhaps you should ask your parrot buddy what to do.
But if you were really serious about this, you wouldn't be going after the strawfolk down in the comments, you'd rebut the paper itself.