Although I guess it is impressive to find consistent results within said micro-benchmark, of course, since that hints at something fundamental about our biology (in the reddit AMA linked elsewhere in the comment the author claims to see no difference for gender, educational background, etc; suggesting age is the main factor. That seems significant).
What 'appears random' is better measurement that matches what people try to do.
People are not good at recognizing randomness, they confuse it with homogeneity. True randomness generates more human recognizable patterns than people think. If you as people to generate random string of 1's and 0's, they avoid long strings of 0's or 1's too much.
So what "appears random" is not at all a good measure of what is "actually random" to put it as simply as possible.
Research goal was to measure cognitive ability and randomness is just measure stick. The actual mathematical complexity is correlated but there is human bias. The bias itself is irrelevant if it's constant. The relevant is how closely subjects can generate strings that appear random and complex (randomness with bias) for humans.
In other words
measure = statistical randomness + bias
Because the bias is almost universal (see the modulating factors in the article) it's not interfering with the thing they try to measure.
So in each case what's the gap between actual linguistic proficiency / randomness, and the appearance thereof? And of what value is measuring these human perceptions rather than the actual facts in each instance (like taking all the results and putting them into a scatter output and seeing if there is actually a pattern in the pseudorandom data, or formally analysing the grammar and spelling in question and verifying that it is technically correct rather than just "english sounding" https://youtu.be/gU4w12oDjn8?t=2m)
Can we draw conclusions about that Italian gentleman's ability to make a song that sounds like English pop music "better" than an English pop music song that is actually technically grammatically correct, and use it to infer that he's got better English skills than the writer of the technically correct song?
And if not, why are we trying to make statements about the ability of some randomness souce not based on any actual measure of true randomness?
That there are no hard constraints/expectations like in measuring the quality of a e.g. software random number generator implementation.
They don't expect to find true randomness in the results, just to measure how much randomness (entropy if you will) those various age groups are capable of producing.