grok is 17%? And that's the lowest, most models are like 80%+?
While hallucination is probably closer to 100% depending on the question. This benchmark makes no sense.
While hallucination is probably closer to 100% depending on the question. This benchmark makes no sense.
But the benchmark didn't ask those questions, and it seems grok is very well at saying it doesn't know the answer otherwise.