I stumbled upon LLM Kryptonite – and no one wants to fix this model-breaking bug
theregister.com
theregister.com
> The chatbot started out fine – for the first few words in its response. Then it descended into a babble-like madness. Which went on and on and on and on and … on. Somehow, it couldn't even stop babbling.
Anyone that's tweaked parameters on local LLMs via e.g. llama.cpp can tell you that it's easy to have a model fall into babbling with the wrong params. I often get this when I try out a new model and restart my server, but my test webUI retains settings from a previous GGUF because I didn't refresh.
But let's be clear: the bug is that the model babbles. But here's the intro that the author justifies with this bug:
> It would appear that the biggest technological innovation since the introduction of the world wide web a generation ago has been productized by a collection of fundamentally unserious people and organizations who appear to have no grasp of what it means to run a software business, nor any desire to implement any of the systems or processes needed to affect that seriousness.
Good grief. I feel like I live in a world where hyperbole is the order of the day, every day.
And what was the exact consequence? Was it just babble, if so why would that matter?
You should also be able to get babble if you set temperature very high and that is to be expected.
Also temp 2.0 will do that with most prompts. Here, I have to take word of the author that the consequence was serious enough.
If I put temperature at 2 and prompt "babble to me about some nonsense" it also goes into different symbols languages and errors out. But that is expected behaviour with temp 2.
>>"We've looked over your report, and what you're reporting appears to be a bug/product suggestion, but does not meet the definition of a security vulnerability."
>That left me wondering whether Microsoft's security team knows enough about LLM internals and prompt attacks to be able to grade a potential security vulnerability. Perhaps – but I got no sense from this response that this was the case.
I'm sympathetic with the author. He doesn't trust the boilerplate response he got. Calling it a "product suggestion" doesn't inspire confidence even if I understand why they'd call it that.
I think that the author doesn't understand what happened, but it must look concerning to them. Maybe it started outputting code as babble because more LLMs have been trained with that? That could be concerning without coding experience, especially if the words that the author can understand sound important or relevant to the LLM itself.
But there are strong and confident claims from the author "What do I do with this powerful and potentially dangerous prompt".
So is the output or result of this prompt anything more than what you would get from high temperature settings?
It's hard for me to think of a security vulnerability here.
Also the example links about prompt attacks etc are about getting past an LLMs censorship. I don't think these are security flaws. The reason why LLMs like these have censorship is just basic PR. I don't think it's a huge issue if it's possible to bypass and ask steps to do something illegal. If anyone seriously wants to do that, they will find a way outside the ChatBot anyway.
Anyone willing to go to Darkweb to get this prompt attack and then use the prompt attack in the actual LLM itself where it might get logged would only put them at risk of getting caught compared to if they just read the instructions on how to do illegal things on Darkweb.
There are other bugs/exploits that are like “if I ask for ‘none beef, no sides, hold the toppings’, the restaurant will serve me an empty box without any food”. These are much less dangerous.
I won’t deny that it’s very cool and novel that it apparently works at all restaurants regardless of what they serve. I’d be fascinated to read an explanation of why that is. Fundamental property of the architecture? Something accidentally present in every data set? A fundamental property of textual data sets themselves?
But based on the apathetic reactions the author has received, it seems like providers universally agree this is a bug of the second kind. One is perhaps reminded of the old tale of the man who goes to the doctor and says “it hurts when I do this” - to which the doctor says “don’t do that, then”.