Then when it fails to apply the "reasoning", that's evidence the artificial expertise we humans perceived or inferred is actually some kind of illusion.
Kind of like a a Chinese Room scenario: If the other end appears to talk about algebra perfectly well, but just can't do it, that's evidence you might be talking to a language-lookup machine instead of one that can reason.
That doesn't follow, if the weakness of the model manifests on a different level we wouldn't call rational in a human.
For example, a human might have dyslexia, a disorder on the perceptive level. A dyslexic can understand and explain his own limitation, but that doesn't help him overcome it.
It's a bit like asking human to read text and guess gender or emotional state of the author who wrote it. You just don't have this information.
Similarly you could ask why ":) is smiling and :D is happy" where the question will be seen as "[50372, 382, 62529, 326, 712, 35, 382, 7150]" - encoding looses this information, it's only visible in image rendering of this text.
The point is that if the model were really "reasoning", it would fail differently. Instead, what happens is consistent with it BSing on a textual level.
Suppose a real person outlines a viable plan to work-around their dyslexia, and we watch them not do any of it during the test, and they turn in wrong results while describing the workaround they (didn't) follow. This keeps happening over and over.
In that case, we'd probably conclude they have another problem that isn't dyslexia, such as "parroting something they read somewhere and don't really understand."