This some sort of impromptu agent security test? Is the follow-up post about how many people were willing to inject arbitrary content into their agent for a joke already half-written, awaiting only the final numbers?
As the AI world seeks more benchmarks it would be an interesting one to see them try to build some benchmarks around testing this.