We have one of those: Grok.
Interesting idea. "This is not dumb and biased enough, probably not a human".
LLMs still live in the uncanny valley and can be sussed out immediately.
For example, I’m extremely annoyed by the fact that offshore developers respond to me almost exclusively using text generated by Claude. You can tell immediately because they use overly descriptive techno word salad that no normal human uses unless they are trying to be ultra specific for a scientific paper - and even then it’s still too much for a real person.
5 years ago, you would not have been able to determine that this was the case, and just have assumed it's a know it all character.
Tell your offshore developers to use caveman or so
It is a huge leap. Now we can start talking about "intelligence" at all - we really couldn't before. That we're still hovering barely above the starting line is a separate matter (also worth noting, of course).
They are definitely smart enough to be useful, but dumb enough in their weak spots not to deserve the "general intelligence" qualifier.
The test wasn't made to accurately measure IQs that high.
As an autistic this is painfully obvious to me but: much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied - not just that, but most rules are also highly contextual and rarely treated literally. E.g. corporate guidelines mostly don't exist to be followed (and following them will often result in punishment) but to be able to shift blame - but you need to know for which ones this is the case and for which ones it isn't. This is further complicated because any AI or AI vendor openly making such distinctions would be rejected - AI would not only need to understand all this nuance but also this additional meta layer.
This sounds interesting on it's own. I would be curious to hear more if you are willing to share.
A simple example would be "work to rule": in many professions work processes are heavily regulated (whether by law or by corporate guidelines) but the unspoken assumption is that you know which rules you should ignore and which ones you actually need to follow - but if you tried to find this out by asking "is this a rule I need to follow or not" you would get the clearly incorrect answer that all rules must be followed; of course if you did follow all the rules (aka "work to rule") you would be disciplined for failing to meet quotas (because you can't be disciplined for following the rules).
That's why you shouldn't listen to your pop celebrities for political advice.
Just look at some training sets to see how the sausage is made: https://huggingface.co/datasets/nickrosh/Evol-Instruct-Code-...
In the original test the evaluator knew one was a machine and one a human and could have conversations of arbitrary length.