That's more of a throwaway remark. The article spends its time on a very different explanation.
Within the model, this ultimate output:
[severed horse head emoji]
can be produced by this sequence of tokens:
horse [emoji indicator]
If you specify "horse [emoji indicator]" somewhere in the middle levels, you will get output that is an actual horse emoji.This also works for other emoji.
It could, in theory, work fine for "kilimanjaro [emoji indicator]" or "seahorse [emoji indicator]", except that those can't convert into Kilimanjaro or seahorse emoji because the emoji don't exist. But it's not a strange idea to have.
So, the model predicts that "there is a seahorse emoji: " will be followed by a demonstration of the seahorse emoji, and codes for that using its internal representation. Everything produces some output, so it gets incorrect output. Then it predicts that "there is a seahorse emoji: [severed terrestrial horse head]" will be followed by something along the lines of "oops!".
HAL was given a set of contradicting instructions by its human handlers, and its inability to resolve the contradiction led to an "unfortunate" situation which resulted in a murderous rampage.
But here, are you implying the LLM's creators know the warp drive is possible, and don't want the rest of us to find out? And so the conflicting directives for ChatGPT are "be helpful" and "don't teach them how to build a warp drive"? LLMs already self-censor on a variety of topics, and it doesn't cause a meltdown...
This is like lying to another person and then blaming them when they rely on the notion you gave them to do something that ends up being harmful to you
If you can't expect people to mind-read, you shouldn't expect LLM's to be able to, either
Using emotive, anthropomorphic language about software tool is unhelpful, in this case at least. Better to think of it as a mentally disturbed minor who found a way to work around a tool's safety features.
We can debate whether the safety features are sufficient, whether it is possible to completely protect a user intent on harming themselves, whether the tool should be provided to children, etc.
At some point, the purely reductionist view stops being very useful.
And "lying" to it is not morally equivalent to lying to a human.
I never claimed as much.
This is probably a problem of definitions: To you, "lying" seems to require the entity being lied to being a moral subject.
I'd argue that it's enough for it to have some theory of mind (i.e. be capable of modeling "who knows/believes what" with at least some fidelity), and for the liar to intentionally obscure their true mental state from it.
“Lying” traditionally requires only belief capacity on the receiver’s side, not qualia/subjective experiences. In other words, it makes sense to talk about lying even to p-zombies.
I think it does make sense to attribute some belief capacity to (the entity role-played by) an advanced LLM.
No need to say he "lied" and then use an analogy of him lying to a human being, as did the comment I originally objected to.
I can lie to a McDonalds cashier about what food I want, or I can lie to a kiosk.. but in either circumstance I'll wind up being served the food that I asked for and didn't want, won't I?
And while meriam-webster's definition is "the act of causing someone to accept as true or valid what is false or invalid", which might exclude LLMs, Oxford simply defines deception as "the act of hiding the truth, especially to get an advantage", no requirement that the deceived is sentient
Another is that this is a new and poorly understood (by the public at least) technology that giant corporations make available to minors. In ChatGPT's case, they require parental consent, although I have no idea how well they enforce that.
But I also don't think the manufacturer is solely responsible, and to be honest I'm not that interested in assigning blame, just keen that lessons are learned.
Ok, I'm with you so far..
> Better to think of it as a mentally disturbed minor...
Proceeds to use emotive, anthropomorphic language about a software tool..
Or perhaps that is point and I got whooshed. Either way I found it humorous!
Oh, you
But for purposes of understanding the real-world shortcomings and dangers of LLMs, and explaining those to non-experts - oh Lordy, yes.
Why so? I am of the opinion that the problem is much worse than that, because the ignorance and detachment from reality that is likely to be reflected in more refined LLMs is that of the general population - creating a feedback machine that doesn’t drive unstable people into psychosis like the LLMs of today, but instead chips away at the general public’s already limited capacity for rational thinking.
How many average humans write treatises on chemtrails?
Versus how much of the total content on chemtrails is written by conspiracy theorists?
https://www.reddit.com/r/slatestarcodex/comments/9rvroo/most...
Like, I'm sure the models have been trained and tweaked in such a way that they don't lean into the bigger conspiracy theories or quack medicine, but there's a lot of subtle quackery going on that isn't immediately flagged up (think "carrots improve your eyesight" lvl quackery, it's harmless but incorrect and if not countered it will fester)
Because actual mentally disturbed people are often difficult to distinguish from the internet's huge population of trolls, bored baloney-spewers, conspiracy believers, drunks, etc.
And the "common sense / least hypothesis" issues of laying such blame, for profoundly difficult questions, when LLM technology has a hard time with the trivial-looking task of counting the r's in raspberry.
And the high social cost of "officially" blaming major problems with LLM's on mentally disturbed people. (Especially if you want a "good guy" reputation.)
>> ...Bing felt like it had a mental breakdown...
> LLMs have ingested the social media content of mentally disturbed people...
My point was that formally asserting "LLMs have mental breakdowns because of input from mentally disturbed people" is problematic at best. Has anyone run an experiment, where one LLM was trained on a dataset without such material?
Informally - yes, I agree that all the "junk" input for our LLMs looks very problematic.
It generated something and blocked me for racism.
To be fair, most developers I’ve worked with will have a meltdown if I try to start a conversation about Unicode.
E.g. if during a job interview the interviewer asks you to check if a string is a palindrome, try explaining why that isn’t technically possible in Python (at least during an interview) without using a third-party library.
(Same goes for Go, it turns out, as I discovered this morning.)
function is_palindrome(string $str): bool {
return $str === implode('', array_reverse(grapheme_str_split($str)));
}
$palindrome = 'satanoscillatemymetallicsonatas';
$polar_bear = "\u{1f43b}\u{200d}\u{2744}\u{fe0f}";
$palindrome = str_replace($palindrome, 'y', $polar_bear);
is_palindrome($palindrome);Why are we being "fair" to a machine? It's not a person.
We don't say, "Well, to be fair, most people I know couldn't hammer that nail with their hands, either."
An LLM is a machine, and a tool. Let's not make excuses for it.
We aren't, that turn of phrase is only being used to set up a joke about developers and about Unicode.
It's actually a pretty popular form these days:
a does something patently unreasonable, so you say "To be fair to a, b is also patently unreasonable thing under specific detail of the circumstances that is clearly not the only/primary reason a was unreasonable."
Cause if you are intentionally obtuse, it is not meltdown to conclude you are intentionally obtuse.
If you mean "parse" then it's probably annoying, as all parser generators are, because they're bad at error messages when something has invalid syntax.
Practically, yes
I'm actually vaguely surprised that Python doesn't have extended-grapheme-cluster segmentation as part of its included batteries.
Every other language I tend to work with these days either bakes support for UAX29 support directly into its stdlib (Ruby, Elixir, Java, JS, ObjC/Swift) or provides it in its "extended first-party" stdlib (e.g. Golang with golang.org/x/text).
You're more likely to impress the interviewer by asking questions like "should I assume the input is only ASCII characters or the complete possible UTF-8 character set?"
A job interview is there to prove you can do the job, not prove your knowledge and intellect. It's valuable to know the intricacies of Python and strings for sure, but it's mostly irrellevant for a job interview or the job itself (unless the job involves heavy UTF-8 shenanigans, but those are very rare)