All evals we have are just far too easy! <1% difference is just noise/bad data
We need to figure out how to measure intelligence that is greater than human.
We need to figure out how to measure intelligence that is greater than human.
Math problems being one of them, if only LLMs were good at pure math. Another possibility is graph problems. Haven't tested this much though.