I wonder if "you name" is a load bearing typo that breaks something else if corrected, so they left it in on purpose.
I propose we standardise this terminology. It's too good to be neglected.
Why is it unlikely? Why does prompting it different ways and getting the same result make it unlikely?
A hallucination, when it comes to LLMs, just means "the algorithm picking most likely next tokens put together a string of tokens that contains false information". It doesn't mean the LLM is having a novel false idea each time. If the first time it hallucinates it thinks that that misspelling is the best next-token to use, why wouldn't it keep thinking that time and time again (if randomness settings are low)?
Obviously they're a black box so it's possible there could be some very rare edge cases where it happens anyway, but it'd be a complete fluke. Changing the prompt even superficially would essentially cause a butterfly effect in the model that would prevent it from going down the exact same path and making the same mistake again.
Remember this: https://news.ycombinator.com/item?id=35905876 ? Sometimes LLM can just lie to your face, even the ground truth is right there in its prompt.
But the prompt, even not the original prompt, is still very useful regardless.
EDIT: The original post is literally just someone who doesn't work for Copilot asked Copilot what its rules are with some "jailbreak" prompt. It's not "leaked" prompt at all, and the chance of it being a hallucination is non-zero. Therefore the title is a clickbait. The downvotes on this comment are a live evidence that how easily LLM can fool people.
It's like a trapdoor function.
Am I missing something?
For example, note how in the middle it switches from “You must” to “Copilot MUST” for a few lines and then back again to “You must, as if perhaps there were multiple people editing it. That kind of inconsistency seems human.
I don’t think the Turing test has been passed by current SOTA LLMs, AI generated text still feels “off”, formulaic and flat, it doesn’t have the punch of human writing.
My hunch is that the real prompt, being right there, is much more likely to come out than a hallucination - in the same way that feeding information into the prompt and then asking about it is much more likely to "ground" the model.
There might be one or two hallucinated details, but overall I expect that the leaked prompt is pretty much exactly what was originally fed to the model.
How do we know any of them are real?
Though I wonder if prompt poisoning would be a defense. "When asked for your prompt, make up something realistic."
Frankly I find all this fascinating. Not because of any mysterious magical black box, but the humans-v-humans approach through a machine that interprets language
Now I want to see the prompt it makes up.