To be fair to the authors they don't actually say it is. But then they contrast it with the "0.00% prompt injection attack success rate".
The upshot is kinda the same - this is still evidence that we should be sandboxing our agents. But it doesn't actually challenge Anthropic's "our models are too clever to prompt-inject" vibe.