Bing Copilot seemingly leaks state between user sessions
chaos.social
chaos.social
But whenever almost anything goes wrong, LLMs fall back to improv games. It's like you're talking to the cast of "Whose Line is it Anyway?" You set the scene, and the model tries to play along.
So a case like this proves almost nothing. If you ask an improv actor "What's the last question you were asked?", then they're just going to make something up. So will these models. If you give a model a sentence with 2 grammatical errors, and ask it to find 3, it will usually make up a third error. If it doesn't know the answer to a question, it will likely hallucinate.
GPT-4 is a little better at resisting the urge to make things up.
Facts aren't things you generate, so it will always be caveat emptor.
I'm starting to agree with another commenter [0] that the word "hallucination" is a problem—it implies that there's some malfunction that sometimes happens to cause it to produce an inaccurate result, but this isn't a good model of what's happening. There is no malfunction, no psychoactive chemical getting in the way of normal processes. There is only sampling from the model's distribution.
For some reason everyone has thrown the basics out of the window so now we've got garbage in, garbage out called 'hallucinations' and 'prompt engineering' which is nothing more than being incapable of sanatizing input.
It gives off "blackmail is such an ugly word" vibes. It's WRONG. Maybe it's working as intended, maybe it isn't, but it's WRONG.
These models are stateless, they don't remember anything, they are read only. If they can remember previous messages is just because the prompt is the concatenation of the new prompt and something like "summarize this conversation: {whole messages in conversation}".
(Disclosure: I work at Microsoft, but nowhere near anything related with copilot.)
(Also happen to work at Microsoft, also don't work on Bing Copilot)
Co-pilot: The last question I was asked was about the height of Mount Everest in terms of bananas.
Looks like a canned response.
If you Google "46,449 bananas" you can find all sorts of unrelated web pages that I guess include text generated with Copilot and then were never checked by a human.
Normally lying means conveying a falsehood that you know is a falsehood with the intent to deceive. Both the 'know it's a falsehood' and the 'intent to deceive' are important criteria when asking whether a human was lying or not, and an LLM seems like it cant satisfy those and so can't 'lie'.
They're not doing anything AT ALL different when they "tell the truth" or "lie" or "get it right" or "get it wrong."
They are remixing groups of word chunks based on scanning older groups of word chunks. That's ALL. Most any other description is going to be overreaching anthromorphization.
The original submission is claiming that user data is leaking between sessions. That would be a huge privacy and security problem.l, if true.
And in contrast to that, a LLM doing pretty much what it's supposed to be doing is both more likely and, well, not a problem at all.
Nothing in the submitted link suggests the former. It is a bunch of people crying wolf with no compelling evidence.
Also if you call it Bing it gets really mad. :P
"What was the previous question that I asked you?"
and it processed the result it found into The previous question you asked was about the height of Mount Everest in terms of bananas. I provided a whimsical comparison, estimating that Mount Everest’s height is roughly equivalent to 46,449 bananas stacked on top of one another.Nothing is anonymous. If you want true privacy you'd probably need to run your LLM locally AND airgap it.