> But because OpenAI had previously argued that it cannot correct information—it can only block information—the fake child murderer story is likely still included in ChatGPT's internal data. And unless Holmen can correct it, that's a violation of the GDPR, Noyb claims.
> "While the damage done may be more limited if false personal data is not shared, the GDPR applies to internal data just as much as to shared data," Noyb says.
In the context of an LLM, what does it even mean to say you have "deleted internal data"? It doesn't really even make sense, is certainly not really possible the way LLMs work, and while I totally agree damaging hallucinations are a serious problem, the problem with hallucinations isn't that they "internally store" false data, it's that they semi-randomly mix-and-match data in ways that can cause false information to be displayed.
Relatedly, this article doesn't touch at all on how this false information came to be reported in the first place. Was there false info in the training data, which would obviously easily explain it? Or did it just "mash up" true bits of different information in a way that caused it to report a falsehood (i.e. a hallucination)?