It's autocomplete. Take it outside the region of validity, and all it's got to work with is whatever extrapolation algorithm happens to have emerged from the weights of the model. (https://xkcd.com/2048/ panel 'House of Cards' comes to mind.) To distinguish between humans and LLMs, we don't even have to take advantage of that: we just have to take the context into the region where a naïve extrapolation of written human output diverges strongly from how humans actually respond. (There are ways to defeat this technique, not that anyone uses them.)
From a technical perspective, the fact an LLM does sometimes (appear to) follow instructions is more of a coincidence that then fact it sometimes doesn't.