Embeddings, vectors, and arithmetic
montyanderson.net
montyanderson.net
There is a zone of illegal thoughts, that becomes definable by model-training. A physical boundary in n-dimensional concept-space. An "aligned" or "safe" AI system knows where this boundary is and does not reach inside it. Vectors (embeddings) that would probe it should instead intersect the surface like a ray-trace in graphics, and return the embedded concept at minimum distance to the safe-idea-boundary.
Intuitively, we all know what this zone is. It's the difference between being a wild barbarian and a gentleman. Or being chill vs antisocial. Seeing it in pure math is pretty awesome.
I'm taking issue with this comment because the way it's phrased (illegal thoughts) is worrisome. It brings up the idea that AI alignment is going down the path of controlling human thought, determined of course by certain corporate and political interests.
I also recommend a video by Yannic Kilcher [2], explaining the paper.
This is why we can’t have nice AI things.
After my experience with RAG across a dozen models and god knows how many experiments against parts of Libgen’s archive in topics I’m familiar with, I’m not sure embeddings are actually useful for anything requiring any kind of accuracy. They’re great for low stakes purposes or as a step in a human driven workflow but like LLMs they’re a very fuzzy and often times inaccurate tool.
Embeddings mean that when we have a thought police they can now be more targeted and effective than before. Any thought you express can be objectively measured "using the euclidean distance or cosine similarity" for illegal concepts, and censored, corrected or punished accordingly. I imagine that this will come early for comment sections on the web.
doesn't it infringe my rights if people use my site and my money to harass people or to spread stuff that is against my economic interest, religion, values, etc. or that I just don't like and didn't intend?
if people are going to run around using AI to spread deepfake ragebait memes, shouldn't I get to enforce policies using the same technology they use to pollute the space?
deciding when free speech crosses a line into a nexus to criminal activity like fraud, libel, conspiracy, incitement, disturbing the peace is hard. deciding when it crosses a line into simply violating community standards designed to further the goals of the community and forum owners, and violates the right or expectation of others for civil discourse is also hard.
it's a bias-variance problem, you can ban the n-word but people find other ways to make people unwelcome, a demagogue can say they are just asking questions when they are pretty plainly trying to start a riot and get people harassed and swatted and feel like they will suffer violence if they disagree.
I don't see anything inherently wrong with using AI tools to help make decisions like that. all that AI that controls what shows up in your feed is already implicitly doing that. inherently when one guy thinks he's making a stand for free speech, someone else is going to think he's coddling extremists who are strategically flooding the zone with bullshit, driving out civil discourse and worse.
I'd honestly love to see AI fending off the eternal winter on message boards.
Is question redundant and basic? Direct user to a specialized AI that can explain the topic well and save the regulars from having to have the same discussion as two days ago.
But it is certainly the case that stronger machine text understanding can be used for censorship and oppression. As pretty much all powerful, general tools can be used for nefarious purposes. But it can also be used for a wide range of great purposes as well.