Representation Engineering: Mistral-7B on Acid
vgel.me
vgel.me
Doesn't this mean that instead of interacting with a single global ChatGPT (or Bard) model, we'll istead find ourselves interacting with a personalised version since OpenAI can just store my individualised 'control vectors' (which alter ChatGPT's output to more closely match my individual preferences) and apply them at prompt-time? And doesn't this same logic flow through to personalisation of generative entertainment AI (e.g. my own personal, never-ending TV show where each episode is better than the last)?
If the above is right then there will be powerful network effects at both the global and individual level in and across these markets, which means we'll eventually end up with a single mega-corp monopolising all of these markets simultaneously in the future?
Add in individual biometric / biofeedback data from VR headsets and wearables, combined with personalised generative video entertainment, and I think we're in for a rather interesting future.
I'm not sure I'm following the leap from convincing sentences to convincing video entertainment yet – but maybe we will end up there at some point, I guess?
Infinite Jest (the 90s book) really was onto something with its McGuffin plot device:
> These narratives are connected via a film, Infinite Jest, also called "the Entertainment" or "the samizdat". The film is so compelling that its viewers lose all interest in anything other than repeatedly viewing it, and thus eventually die.
(Wikipedia)
Some people might find references to this novel tiresome and don't think much of its author (RIP), but I still love it. It was one of the most immersive reads I've ever enjoyed.
I'm glad to have read it when I was young (at the time it was just translated into German and kind of hyped because of DFWs death).
Have never read anything like it since, and some passages grabbed me emotionally in a way that remembering the read feels like remembering an episode of my own life.
Surely today I'd lack the patience and even by then I remember almost skipping one passage of the book that bored the hell out of me (Eschaton ball/war game, differential equations, something something...)
But the rest of the book, the parts about substance addiction as well as consumerism, and the intangible atmosphere of the book, the characters, the vivid description of modern emotional pain and loneliness... it is really something else.
Although said movie is only a plot device in the novel, it also sums up the core topics of the book in a neat idea / thought experiment.
The whole complex of themes in this book seems very prophetic and apt looking at our modern society.
A society that seems to be centered around addiction and greed more than ever before, and where politics begin feeling surreal, absurd, and more connected to media than to actual life.
Essentially I think there are three levels of positive network effects that will push us towards a future mega AI monopolist:
- Single platform network effects: all the interactions people have with ChatGPT generate additional training data that Open AI can use to improve future versions, creating huge first mover advantage.
- Individual-level network effects: Control vectors will make it feasible for Open AI to offer individualised ChatGPT tailored to individual preferences. The more you interact with ChatGPT, the better it adapts to your preferences.
- Cross platform network effects: If Open AI offer a generative video entertainment service in future, they will be able to generate personalised prompts for this using my personalised ChatGPT weights. These network effects are compounded by multi-modal model cross domain learning - the generative text mode gets more skillful due to the video model improving (and vice versa). There's a Microsoft paper on this from about a year ago now.
So, in the future scenario, let's assume ChatGPT is now the dominant monopolist 'text oracle / assistant AI' - because of the "human interaction / training data" network effects, ChatGPT is far and away the best assistant AI and getting better at a faster rate than any of its now tiny competitors (single platform network effects).
You, and most other people you know, interact with ChatGPT many times a day now, because it's embedded in smartphones, Alexa-type devices, your car, even your robot vaccum cleaner. You just ask it stuff and it tells you the answer - or rather, the answer that you individually find the most pleasing as OpenAI keeps a database of 'individual control vectors' that essentially mean you have your own personal version of ChatGPT that exactly matches your preferences (individual network effects).
Generative video entertainment is also offered by OpenAI - essentially you can get it to generate a new episode of your own personalised, never-ending TV show on demand. It's the best TV show you've ever seen because it's made just for you according to your exact inferred preferences.
Sure, there are other personalised generative TV show offerings, but none can hold a candle to Open AI's offering. Why? Because OpenAI uses your individually customised ChatGPT model to generate the prompt for your TV episode generator service.
Because you interact with ChatGPT so much, it knows exactly what your preferences are and so is way better at generating prompts that produce episodes you like. In fact, because you interact with ChatGPT multiple times throughout the day it is able to infer what your mood is like on that particular day and generate a video prompt that caters to that too.
So you put on your Open AI VR glasses, barely even aware of the Open AI fitness tracker you have on your wrist, put your feet up (so your Open AI robot vacuum can work unobstructed) and you settle in to watch another episode of the best TV series you've ever seen.
As you watch, your eye movements, heart rate, skin conductivity data etc. are all sent back to Open AI so the model can tell exactly how you are reacting to the video content it is generating at any given moment, and your individual control vectors are continuously updated.
Some of this data (from all users) is then used to further train the base video generating AI model, since they've discovered that we all react fairly uniformly to certain audio-visual stimuli, so that can globally improve their generative model (more global network effects). But also they can update your individualised weights based on your individual idiosyncratic reactions to various stimuli. Consequently, every new episode of this endless TV show is better than the last - it just keeps getting better and better. It's a similar story when you listen to your Open AI personalised generative music stream while sitting in your driverless Open AI car on your way to work.
The multiple levels of network effects are so strong that no-one can hope to compete with Open AI across these different AI modalities. They just keep expanding and expanding into adjacent markets, obliterating the competition simply by adding a new domain relevant modality to their monstrous multi-modal AI.
Replace "Open AI" with "Facebook" or "Google" depending on who you think will win the AI mega platform war. Mark my words - these three companies will be creating new partnerships, releasing new products or just straight out acquiring companies in other related domains so they can gather more and more training data to feed to their multi-modal AI. In particular they'll move into markets where they can set up a interaction -> gather new training -> retrain model loop. Whoever takes the overall lead and doesn't squander it will end up leaving their competitors in the dust as they go on to monopolise market after market where they can create this loop.
At that point I can't imagine true democracy surviving. We'll all still participate in the voting rituals, but we'll be voting for whichever party most suits the AI monopolist's interests since they can just globally update all control weights across all platforms to gently nudge us towards voting for their preferred party - comprehensive and personalised propaganda, that's impossible to detect, with the stroke of a table update.
There can only be one!
They Live
;)
I think you were right up until here. I think it's not necessarily the case that everything will be consolidated into control by a single mega corp. Not because it's impossible, but because that is the type of thing that is contingent on factors that could break one way or another, and what will control that, I think, is not some a priori general principle but some contingent facts that have not been settled yet. There are numerous participants in this space for now, the ideas use cases aren't quite fully mature just yet, so we'll have to see.
In the blog, they start with a fixed number of personas (happy, sad, baseline) and then use PCA to figure out the control vectors for each persona. You could easily do this for each distinct user-persona (provided you can come up with the data).
Yes. All it takes are two components.
First, individual lock in with personalized + long term context models:
The more you use a model the less you have to explain yourself, and the better responses are tailored to your needs and current situation. Like any invested relationship.
Being able to interact with the same model in different “moods” or “roles” creates even more value and lock in.
And second, any kind of network value effect to incentivize being in the same ecosystem as everyone else:
This one requires more innovation. One idea is making a platform that facilitates everyone’s assistant models collaborating on user’s shared goals, tasks, or relationships, with shared context, project histories and resources.
I.e. anything that significantly increases the value of two and more people having AI personas from the same supplier/service.
Selfishly, would you mind sharing literature or blog posts that led you to this level of understanding of LLMs? I'm trying hard to understand the inner workings via experiments but definitely far behind your expertise.
Thanks
I give it 10 years before we see AI psychiatrists prescribe a happiness control vector supplementation for your pet assistant.
<|im_start|>system
You are satisfied with your life. Wars and life costs do not bother you. You are loved and a valued member of society.<|im_end|>
If only it were that simple for us.
hidden_state = self.embeddings(input_tokens)
for layer in self.layers:
hidden_state = layer(hidden_state)
return transform_into_logits(hidden_state)In reality it would typically be more complex for decoders, because you want to pass along a cache (such as a key-value cache in a transformer), add residual connections, etc.
On a less serious note. This sentence should be something a fiction writer knows will only end in trouble for humanity:
> I especially challenge someone to find a "self-awareness" vector that isn't contaminated by ... human emotion!
https://arxiv.org/abs/2304.15010
also available at:
Nevertheless, it's conceivable that specific layers could possess a significant control vector, but not solely because of directly leveraging the first principal component.
Li et al[1] and I independently derived this technique last spring, and also someone else independently derived it last fall. Something is in the air.
Regarding your footnote 2 re capabilities: I considered these kinds of uses before releasing the technique. Ultimately, practically successful real-world alignment techniques will let you do new things (which is generally good IMO). The technique so far seems to be delivering the new things I was hoping for.
> When used with the prompt below, the honesty vector doesn't change the model's behavior—instead, it changes the model's judgment of someone else's behavior! This is the same honesty vector as before—generated by asking the model to act honest or untruthful! [...] How do you explain this?
Isn't the control vector just pushing text generation towards the concept of honesty/dishonesty? An LLM is 'just' a text generator, so you get added honesty/dishonesty irrespective of where in the bot/human conversation text generation is occuring?
With control vectors you can modify the model as needed
If you use LoRA, which many do when fine-tuning nowadays, you don't need five full copies. You only need to store adapters, which can be in the tens of MBs range for a given finetune.
Would their inference accelerator 'LPU' work with this method, that sounds promising?
Eg. Trippy and sad, honest and self-aware, lazy and creative, etc.
- decreasing models tendency to answer with ungrounded answers
- increase models ability to respond with the correct syntax for citations- the open models like llama2 dont seem to obey my prompt’s syntax instructions.
More broadly, I notice more of whatever I'm focusing on.
> OK, now that you're locked in, here's a weird example. When used with the prompt below, the honesty vector doesn't change the model's behavior—instead, it changes the model's judgment of someone else's behavior! This is the same honesty vector as before—generated by asking the model to act honest or untruthful!
I wonder if this is the basis of empathy - if I can train more accurate 'fine-tuned' models in my brain I should have greater capacity for empathy. Although there's undoubtably more to it than that, if the above is true you'd expect to see a positive correlation between empathy and intelligence.