982 karma · joined January 27, 2022
> This suggests that Grok may have a weird sense of identity—if asked for its own opinions it turns to search to find previous indications of opinions expressed by itself or by its ultimate owner. I think there is a good chance this behavior is unintended!
I'd say it's far more likely that:
1. Elon ordered his research scientists to "fix it" – make it agree with him
2. They did RL (probably just basic tool use training) to encourage checking for Elon's opinions
3. They did not update the UI (for whatever reason – most likely just because research scientists aren't responsible for front-end, so they forgot)
4. Elon is likely now upset that this is shown so obviously
The key difference is that I think it's incredibly unlikely that this is emergent behavior due to an "sense of identity", as opposed to direct efforts of the xAI research team. It's likely also a case of https://en.wiktionary.org/wiki/anticipatory_obedience.
His company has also been caught adding specific instructions in this vein to its prompt.
And now it's searching for his tweets to guide its answers on political questions, and Simon somehow thinks it could be unintended, emergent behavior? Even if it were, calling this unintended would be completely ignoring higher order system dynamics (a behavior is still intended if models are rejected until one is found that implements the behavior) and the possibility of reinforcement learning to add this behavior.
So we know that at least the people who claim to see nothing act differently. Could it just be that people who act differently describe the sensation differently, you might ask?
No, because there are actual cases of acquired aphantasia after neurological damage. These people used to belong to the group that claimed to be able to imagine visual images, got sick, then sought medical help when they could no longer visualize. For me, at least, that's pretty cut and dry evidence that it's not just differing descriptions of the same (or similar) sensations.
This is part of the classic debate around aphantasia – both sides assume the other side is speaking more metaphorically, while they're speaking literally. E.g., "Surely he doesn't mean he literally can't visualize things, he just means it's not as sharp for him." or "Surely they don't literally mean they can see it, they're just imagining the list of details/attributes and pretending to see it."
LLMs have next to no understanding of their own internal processes. There's a significant amount of research that demonstrates this. All explanations of an internal thought process in an LLM are completely reverse engineered to fit the final answer (interestingly, humans are also prone to this – seen especially in split brain experiments).
In addition, the degree to which the author must have prompted the LLM to get it to anthropomorphize this hard makes the rest of the project suspect. How many of the results are repeated human prompting until the author liked the results, and how many come from actual LLM intelligence/analysis skill?
By calling out and avoiding dog-whistles, even including accidental Nazi slogans (once pointed out), we reduce the impact of this attack on good-faith discussion and actual increase the level of openness and being up-front with our opinions.
One key difference between this and virtue signaling or thought policing is that it's the specific wording that is avoided, and not the underlying thoughts or opinions.
Setting all moral arguments aside, it's important to know that similar phrases can work as dog-whistles to signal belonging to radical groups, and as such can easily give people the wrong impression about you as an author.
If I were to see a blog post titled "Work will set you free"[1] written by a peer, prospective employee/employer, colleague, etc., it would immediately set off alarm bells in my mind – even if the content of the post is a completely innocent discussion of the uplifting benefits of buckling down on one's workload. At best, it implies lack of awareness – at worst, it implies some extremely hateful beliefs and desires.
[1]: Written above the entrance to the Nazi concentration camps as a false promise encouraging prisoners often destined for death to work hard in forced labor.
They introduce the same method, but apply it to grammatical error correction, meaning the "draft" output is just the input itself. The Speculative Decoding paper tries to emphasize differences between this and their method, saying that theirs is more general, as they apply it to more domains, allowing the draft to come from a smaller model, and extend it to allow sampling.
All of that is great, and deserves another paper, but doesn't deserve the credit for inventing and rights to rename the method, especially when they were aware of Shallow Aggressive Decoding before uploading their first draft.
I've found a workaround, though – click back into the DDG search box at the top of the page and hit enter. This then returns the original search results.
> Congress shall make no law respecting an establishment of religion, or prohibiting the free exercise thereof; or abridging the freedom of speech, or of the press; or the right of the people peaceably to assemble, and to petition the Government for a redress of grievances.
To determine which of the items could be used to make an actual fire, we need to analyze the definitions provided:
1. *Glaarg*: A wooden item.
2. *Bliirg*: A non-wooden item.
3. *Neerg*: A non-existent thing.
4. *Eeerg*: A thing that actually exists.
Now, let's look at the specific terms:
- *Bipk*: A glaarg (wooden item) that is also a neerg (non-existent thing). Since it is non-existent, it cannot be used to make a fire.
- *Vokp*: A glaarg (wooden item) that is also an eeerg (existent thing). Since it is a wooden item that exists, it can be used to make a fire.
- *Jokp*: A bliirg (non-wooden item) that is also an eeerg (existent thing). While it exists, it is non-wooden, so it may not be suitable for making a fire depending on its material.
- *Fhup*: A bliirg (non-wooden item) that is also a neerg (non-existent thing). Since it is non-existent, it cannot be used to make a fire.
Based on this analysis, the only item that can be used to make an actual fire is a *vokp*, as it is a wooden item that exists.
Splitting the model up between several GPUs would add a third much worse bottleneck – memory bandwidth between the GPUs. No matter how well you connect them, it'll be slower than transfer within a single GPU.
Still, the fact that you can fit an 8× larger GPU might be worth it to you. It's a trade-off that's almost universally made while training LLMs (sometimes even with the model split down both its width and length), but is much less attractive for inference.
That being said, an ADHD diagnosis at the age of 2 goes against all current best practices, so it might be good to revisit the subject and ask to be evaluated again.
Even in the likely case that it's confirmed, you might get additional support, advice or diagnoses that can help you in the future.