That core point is like a variation on a lemma in a paper I recently published (Lemma 2 in [1]): There can be no single deterministic interactive reward-giving environment which can, by its lonesome, serve as a good proxy for general intelligence of deterministic agents.
The proof is fairly simple. Suppose E were such an environment. I claim that every intelligent agent is exactly as intelligent (according to intelligence as measured by performance in E) as some "blind" agent, where by "blind" I mean an agent which totally ignores its surroundings.
Let A be any deterministic agent. If we were to place A in the environment E, then A would take certain actions, call them a1, a2, a3, ..., based on A's interaction with E.
Now define a new agent B as follows. B totally ignores everything and instead just blindly takes actions a1, a2, a3, ...
By construction, B acts exactly the same as A within E, therefore, as measured by E-performance, A and B are equally intelligent. So, for any particular environment, every agent is just as intelligent as some "blind" agent. But that's clearly a very bad property for an alleged intelligence measure to possess!
[1] https://philpapers.org/archive/ALEIVU.pdf