Seems like a good benchmark for AGI. Start with things that are easy for humans but hard for LLMs currently.
Ask it to count using a coding tool, and it will always give you the right answer. Just as humans use tools to overcome their limits, LLMs should do the same.