That's what makes it a fair evaluation of its limits
Though I suppose we're testing their model + agent harness here as well. It really _should_ have all of those tools/reasoning available to accomplish a task like the above without issue.
The point is what are the typical use cases for the tool / what are the agreed upon areas of application?
Making the LLM do math with large numbers, I would argue, is not in its typical use case, thought it's at the border.
Asking an image generator model to calculate numbers before running an image sounds definitely NOT like a reasonable use case (do people need it? Will people try using it for this purpose?)