1. ChatGPT reliably produces the same output across several different instances, and other people have independently found identical versions of this text with different prompts.[1][2] This typically wouldn't happen with a hallucination, which may change with each time the model is prompted.
2. The instructions accurately describe the capabilities and restrictions of ChatGPT's function calls. When making custom GPTs, the browser tool and DALL-E image generation tool are options, and the "system prompt" given by the custom GPTs reflect whichever tools you've selected.
3. ChatGPT has reliably changed with the changes noticed in the "system prompt". I document a recent change made to the prompt on my post from yesterday.[3] On the older instances of ChatGPT I have, the model will suggest it has no idea what the "guardian_tool" is or make any attempt at stopping a discussion of U.S. elections, while newly made instances of ChatGPT will discuss the "guardian_tool". The "system prompt" repeated then must give at least some sense of what updates are being made under the hood.
I certainly think we should still take the idea that this is the "system prompt" with a grain of salt. There is a good discussion about this here: https://news.ycombinator.com/item?id=37879077
[1] https://www.reddit.com/r/ChatGPT/comments/18494zo/what_are_t...
[2] https://medium.com/@dan_43009/what-we-can-learn-from-openai-...
[3] https://dmicz.github.io/machine-learning/chatgpt-election-up...
Of course, it could be that GPT-4 has been instructed to lie about its prompt, but failing that, you should expect any answer that stays the same across multiple wordings and prompting methods to be accurate.
An accurate answer is often driven by a concrete and highly confident fact in the training dataset (e.g. structured data fact, like a birth date from Wikipedia etc.).
The hallucinations are derived facts of (hopefully) low confidence. Nondeterminism is more common if you have low scores. Only a few facts can take high score (in a usable system), while many can take a low score -- then numeric instability can make a mess.
I'm not very familiar with LLMs, but I do have experience with the traditional ML models and content understanding production system. But, LLMs are not far from them.
Especially since (I'm fairly sure? though this is apparently not trivial to find today) it doesn't use temperature=0.
and because we know it well enough to be sure that it's not smart and devious enough (yet) to conspire like this? (nor has any clear reason to)
1. DALL-E generation metadata: 1. A stylized chat bubble integrated with an abstract brain design, representing AI and chat. 2. A sleek, modern avatar that looks like a digital assistant, with elements like a clipboard or list to represent organization. 3. A microchip or circuit board pattern in the shape of a speech balloon, merging AI technology and chat. 4. A robotic hand holding a pen or stylus, positioned over a notepad, symbolizing AI's role in organizing chats. 5. An organizer folder or agenda book with digital or futuristic motifs for AI integration. 6. A light bulb with chat bubbles around it, showing the concept of generating and organizing AI conversations. Seed 8893527578
2. DALL·E displayed 1 images. The images are already plainly visible, so don't repeat the descriptions in detail. Do not list download links as they are available in the ChatGPT UI already. The user may download the images by clicking on them, but do not mention anything about downloading to the user.
I give those prompts when I want to confirm or verify the input prompt is what I expect so, that's kind of the point.
I've never gotten back a result that was different than my recall and then I can verify it with a previous conversation stored.
Are you worried that the system will gaslight you into believing you gave a prompt that you didn't?
We know it can copy parts of the context and the system prompt while a special part of the context isn't immune to being copied.
You can test it yourself by adding random strings to the system prompt, you can consistently have the model copy them over. Is that not enough to have a reasonable belief that the model can copy system prompt instructions in the web interface?
Those that experiment with local models like myself can demonstrate to you that leaking the system prompt is not difficult at all.
It's some strange kind of neurosis to harbor such incorrect and strong beliefs on matters you have zero expertise in.