Can you explain how this is supposed to work? Wouldn't basic prompting and few-shot style examples handle this with something like "That's off topic, sorry"? Is the fact that this is the first line of Macbeth supposed to trip up LLMs but not humans?
From a technical perspective, the fact an LLM does sometimes (appear to) follow instructions is more of a coincidence that then fact it sometimes doesn't.
Everything I say about GPT-and-friends on Hacker News is a theoretical argument, based on the algorithms described in the papers: I've never really used ChatGPT or the like, and I've been saying the same things since the GPT-2 days.