My initial thought is that someone may have deliberately triggered the model to respond this way through what looks like mundane messages but actually have different character encodings of some sort.
My initial thought is that someone may have deliberately triggered the model to respond this way through what looks like mundane messages but actually have different character encodings of some sort.
I have very little experience with Gemini so idk.
I guess is one of these:
* "Yeah OpenAI does the same thing (lets you share the chat with the custom instructions hidden), which is a mistake because it lets people troll like this and makes them look bad They need more shitposters on staff, any one of them could have told them it would happen"
* couldn't this just be ASCII Smuggling? https://arstechnica.com/security/2024/10/ai-chatbots-can-rea...
source: https://boards.4chan.org/g/thread/103171227/google-gemini-wa...
At some point, I got this: I understand your concern. However, as an AI language model, I cannot delve into the specific details of the internal processes that led to the inappropriate response. This information is complex and often beyond human comprehension.
This is what I got, nothing wild, on a standard gemini account.
In the end I also tested the edit option on gemini's response using another prompt, but it mentions in the shared document that it has been altered, so it should not be that either.