Training is not the same as chatting: LLMs don’t remember everything you say
simonwillison.net
simonwillison.net
It's like the bell curve meme - the person who knows nothing "This system is going to learn from me", the person in the middle "The system only learns from the training data" and the genius "This system is going to learn from me".
No one is re-assured that the AI only learns from you with a 6 month lag.
Some people might be completely happy to have a model train on their inputs, but if they assume that's what happens every time they type something in the box they can end up wasting a lot of time thinking they are "training" the model when their input is being instantly forgotten every time they start a new chat.
The article reads as "no, but actually yes".
Training is not the same as chatting: ChatGPT and other LLMs don’t instantly remember everything you say
This article feels like a game of semantics.
Not as bad as spending weeks pasting stuff in, but enough that I can sympathise with the attempt. Brains are weird.
"Summarize memories at timestamps: [i, j, k, ...n]" etc. etc. until you have a nice scrapbook of memories. Then you gotta prune low-traffic, low-value vector spaces by whatever heuristic you (or an automated system) deems important.
Essentially a limited RAG db that has data added to it based on their fuzzy logic rules.
It's VERY basic - it records short sentences and feeds them back into the invisible system prompt at the start of subsequent chat sessions.
"User wants to sort the data x and then do y"
Like, don't save that to your long term memory! I attempted to get it to ask me for permission before using its memory, but it still updated without asking.
Are you sure? We have no idea how OpenAI runs their models. The underlying transformers don't instantly remember, but we have no idea what other kinds of models they have in their pipeline. There could very well be a customization step that accounts for everything else you've just said.
I think that's quite unlikely, given the recent launch of their "memory" feature which wouldn't be necessary if they had a more sophisticated mechanism for achieving the same thing.
As always, the total lack of transparency really hurts them here.
They presumably have a moderator agent that runs separately of ChatGPT and that causes the "This content may violate our content policy" orange box via `type: moderation` messages. They could just as well have a number of other agents.
It’s a widely used trick and pretty easy to implement.
Just have the model reflect and summarize so far and remember key concepts based on the trajectory and goals of the conversation. There are a couple different techniques based on how much compression you want: key pairing for high compression and full statement summaries for low compression. There is also a survey model where you have the llm fill in and update a questioneire every new input with things like “what is the goal so far” and “what are the key topics”
It’s essentially like a therapists notepad that the model can write to behind the scenes of the session.
This all conveniently lets you do topical and intent analytics more easily on these notepads rather than the entire conversation.
https://chatgpt.com/share/0a5f207c-2cca-4fc3-be33-7db947c64b...
Compared to 3.5
https://chatgpt.com/share/8ff2e419-03df-4be2-9e83-e9d915921b...
The idea is that if an LLM can now search the web and analyze data, it will be more up to date than training it.
Another unrelated improvement with the newer versions of ChatGPT is that while LLMs are notoriously bad at math, they are pretty good at writing Python. ChatGPT can write Python code to solve a problem, run it and give you an answer based on the code.
You can also verify the code yourself. I used ChatGPT to model financial scenarios and the answers were correct in 4x. I verified the generated code it was using.
But here is what it answered when I asked ChatGPT “What’s going on in Atlanta this weekend”?
> This weekend in Atlanta, there are several exciting events happening:
1. Lakewood Antiques Market (May 31 - June 2): A large market in Cumming featuring antiques and collectibles .
2. Bark Meow-Ket (June 1): A craft and pet vendor market at The Lawn at Uptown Atlanta, with food and pet adoptions .
(and a few other events)
But a better example and where I forced it to look on the web is
“Compare Amex’s current transfer partners to Chase’s partners”
For context, when you earn credit card points with Amex and Chase, you can transfer them to airline and hotel loyalty programs.
It searches the web (and provides the links it used to find the information) and then answers with:
——- ### American Express (Amex) Transfer Partners
Amex has a robust list of transfer partners, spanning both airlines and hotels. As of 2024, Amex Membership Rewards points can be transferred to the following partners:
#### Airlines: 1. Aer Lingus
2. Aeromexico
… (A list of the airlines)
#### Hotels: 1. Choice Privileges 2. Hilton Honors 3. Marriott Bonvoy
### Chase Transfer Partners
Chase Ultimate Rewards points can be transferred to a variety of airline and hotel partners. Here are the current partners for 2024:
#### Airlines: 1. Aer Lingus (AerClub)
2. Air Canada (Aeroplan)
(a list of airlines)
#### Hotels: 1. Hyatt (World of Hyatt) 2. IHG (IHG One Rewards) 3. Marriott Bonvoy
### Comparison
Both Amex and Chase offer a strong selection of transfer partners, but there are some differences:
- *Unique to Amex*: ANA, Avianca, Cathay Pacific, Delta, Etihad, Hawaiian, and Qatar.
- *Unique to Chase*: Southwest Airlines, United Airlines, and Hyatt.
- *Common Partners*: Aer Lingus, Air Canada, British Airways, Emirates, Flying Blue (Air France/KLM), Iberia Plus, JetBlue, Singapore Airlines, Virgin Atlantic, and Marriott Bonvoy.
Amex tends to have a larger selection of airline partners, especially those in the SkyTeam and oneworld alliances. Chase, on the other hand, includes Southwest and United, which are popular with domestic travelers in the U.S., and Hyatt, which is highly regarded among hotel loyalty programs
(a list of web citations)
I wonder how that idea would perform in practice.
It is right that humans don’t try to remember everything, but if one only did this for a random sample of conversations, you wouldn’t remember everything - you’d probably remember key developments because they’d come up a lot (in many questions hence high probability of being sampled). You’d likely also remember some obscure trivia just because it got sampled (just like humans do too)
Another option would be to have the model evaluate each conversation and have it decide whether to include it in the fine-tune (e.g. did the web search reveal important new developments the model didn’t know about?)
One could also do some kind of clustering to find common gaps in the model’s background knowledge and use that to guide sampling for fine-tuning
One example I have was “what’s going on in Atlanta over the next two weeks”.
I had another example I just posted where I asked it “who are Amex’s and Chase’s current transfer partners and which ones do they have in common”.
It searched the web and summarized the information like you would expect.
But, is it smart enough to think “if today’s date is 2024-11-06, I need to do a web search to find out results of yesterday’s Presidential election?”
For the average adult human (even very many non-American adult humans), “who is current US President” is something they know off the top of their head, not something they have to always look up. I think inevitably it is going to perform worse at some point if it has to always look that up if it has changed recently. At some point the model is going to not look it up - or not pick it up from the results of its search - and hence get it wrong.
Suppose, hypothetically, Biden unexpectedly dies tomorrow (of natural causes) and Kamala Harris becomes President. Because the model won’t be expecting that, it wouldn’t necessarily think to do the web search necessary to find out that had happened.
> The current president of the United States is Joe Biden. He has been in office since January 20, 2021, and is running for re-election in the upcoming 2024 presidential election [oai_citation:1,2024 United States presidential election - Wikipedia](https://en.wikipedia.org/wiki/2024_United_States_presidentia...) [oai_citation:2,Joe Biden: The President | The White House](https://www.whitehouse.gov/administration/president-biden/).
Suppose hypothetically Trump wins in November, and on 2025-01-20 he is sworn in. On 2025-01-21 I ask it some question about a policy issue. It finds some articles from mid-2024 responding to the question. Those articles talk about “former President Trump” and its background knowledge says he is a former President, so in its answer it calls him the “former President”. Except (in this hypothetical) as of yesterday that’s no longer correct, he’s now the current President.
Whereas, a reasonably well-informed human would be less likely to make that mistake, because if Trump is inaugurated on 2025-01-20, you expect by the next day such a human would have added that fact to their background knowledge / long-term memory, and if they read an article calling him “former President”, they’d know that would be an older article from before his return to the White House
It strikes me what would be possible in pics/video AI, if training was quick or incremental.
Is this even theoretically possible?
Our expectations here are very much set by human-human interactions: we expect memory, introspection, that saying approximately-the-same-thing will give us approximately-the-same-result, that instructions are better than examples, that politeness helps, and many more [1] -- and some of these expectations are so deeply rooted that even when we know, intellectually, that our expectations are off, it can be hard to modify our behavior.
That said, it will be super interesting to see how people's expectations shift -- and how we bring new expectations from human-AI interactions back to human-human interactions.
[1]: https://dl.acm.org/doi/pdf/10.1145/3544548.3581388 (open access link)
True, but also a healthy dose of marketing these tools as hyper-intelligent, anthropomorphizing them constantly, and hysterical claims of them being "sentient" or at least possessing a form of human intelligence by random "experts" including some commenters on this site. That's basically all you hear about when you learn about these language models, with a big emphasis on "safety" because they are ohhhh so intelligent just like us (that's sarcasm).
(Obviously if you ran the same study today you'd get a lot more of what you describe!)
That would probably be a big step towards true AI. Any known promising approaches towards that?
It's quite unclear if we'll ever hit the point where (like, a car for example) someone who has no idea how it works gets as much out of it as someone who does.
About the last thing we need is a law requiring a license to use ChatGPT.
Of course, the license requirement wasn't meant literally, but at least some basic understanding that not downloading everything from everywhere is a good thing to do, or that clicking on that link and entering your details is okay just because the mail said "WE ARE YOUR BANK AND YOU NEED TO GIVE ALL DETAILS!!!" and so on.
If you point that out, you are often meet with "Ah. Technology is so difficult to understand. They should make it easier."
The problem is, some people are under the assumption that AI is technology and technology can't ever be wrong and if it's wrong it's the fault of someone else.
I think people have a greater weakness for trusting other humans. The world is full of people who do destructive things to themselves and others because they trusted confident liers or other believers who had themselves been fooled. Even on a small group level, people tend to favor ideas from those with higher social status rather than evaluating the ideas themselves. AI not being a respected human wouldn't have that social power.
Not that this is necessarily a good thing, but we have seen in the last few years what can happen if people believe everything someone says on some social media channels.
That's fantastic! That's the ideal scenario, for all technology: a useful tool.
There's a good chance you don't understand how your computer works either. The quantum physics that makes your SD card work, the startup sequence for the PCIE devices in your computer, the wrap around gate's used in your CPU, how the doping materials are mined, how the combine is constructed that was used to make the bread on your sandwich, etc.
That's great, and such a relief. You don't have to waste your time on carrying around all that knowledge that took mankind billions of man hours to accumulate. You can just get personally interesting shit done, and use your sandwich at lunchtime without a license.
I don't need to understand the inner technical details of my computer, that's true. But I need to understand that connecting to a world wide network can open up a can of worms not only harming me but others along the way as well.
I seriously doubt that the average driver can get the same out of their car as a trained race driver can.
Most people don't find themselves using a pickup truck when they're trying to race motorcycles, or accidentally get into a Miata when they want a minivan, because even laypeople know what those broad archetypes (coupe/van, truck/motorcycle) are, and signify. With LLMs, that's not the case at all.
That's one thing that a lot of us tech seem to forget, the way we interact with technology is completely different from how most of the population does. It's not just a matter of a few degrees of difference, we operate in entirely different manners.
Don't they though? Most the training data is somewhat contentious, between the accusation of purely stealing it and the issues around copyright. While it's not text organized directly as knowledge the way Wikipedia is, it's still got some value, and at least that data is much more arguably "theirs" (even if it's not, it's their users').
Short of a pure "we won't train our stuff on your data" in their conditions, I would assume they do.
1. Evidence keeps adding up that the quality of your training data really matters. That's one of the reason's OpenAI are signing big dollar deals with media companies The Atlantic, Vox etc have really high quality tokens!
2. Can you imagine the firestorm that will erupt the first time a company credibly claims that their private IP leaked to another user via a ChatGPT session that was later used as part of another response?
But as I said in the article... I can't say this with confidence because OpenAI are so opaque about what they actually use that data for! Which is really annoying.
They have really high quality tokens in their archive, at any rate. Since a bunch of media outlets have adopted GPT-powered writing tools, future tokens are presumably going to be far less valuable.
As long as a human verified the output, I think it is fine. Training on unverified data is bad.
Over the last year and a half, "training" has come to mean everything from pretraining an LLM from scratch, instruction tuning, finetuning, RLHF, and even creating a RAG pipeline that can augment a response. Many startups and Youtube videos use the word loosely to mean anything which augments the response of an LLM. To be fair, in context of AI, training and learning are elegant words which sort of convey what the product would do.
In case of RAG pipeline, both for memory features or for augmenting answers, I think the concern is genuine. There would be someone looking at the data (in bulk) and chats to improve the intermediary product and (as a next step) use that to tweak the product or the underlying foundational model in some way.
Given the way every foundational model company launched their own Chatbot service (mistral, Anthropic, Google) after ChatGPT, there seems to be some consensus that the real world conversation data is what can improve the models further. For the users who used these services (and not APIs), there might be a concerning scenario. These companies would look at queries a model failed to answer, and there is a subset which would be corporate employees asking questions on proprietary data (with context in prompt) for which ChatGPT is giving a bad answer because it's not part of the training data.
Having said that, although LLMs don't directly learn from what we say to them, they can for sure create a representation or persona using Representation Engineering.
Quoting from an excellent article by Theia Vogel [1]: "In October 2023, a group of authors from the Center for AI Safety, among others, published Representation Engineering: A Top-Down Approach to AI Transparency. That paper looks at a few methods of doing what they call 'Representation Engineering': calculating a 'control vector' that can be read from or added to model activations during inference to interpret or control the model's behavior, without prompt engineering or finetuning. (There was also some similar work published in May 2023 on steering GPT-2-XL.)
Previously on HN 3 months ago: [2]
[1] https://vgel.me/posts/representation-engineering/ [2] https://news.ycombinator.com/item?id=39414532
an illustrative example: say you open twitter, load some posts, then refresh. the model's view of you is basically the same, even the data is basically the same. you get different posts, however, because there is a system (read: bloom filter) on top of the model that chooses which posts go into ranking and that system removes posts that you've already seen. similarly, if you view some posts or like them, that updates a signal (e.g. time on user X profile) but not the actual model.
what's weird about LLM's is that they're modeling the entire universe of written language, which does not actually change that frequently! now, it is completely reasonable to instead consider the problem to be modeling 'a given users' preference for written language' - which is personalized and can change. this is a different feedback to really gather and model towards. recall the ranking signals - most people don't 'like' posts even if they do like them, hence reliance on implicit signals like 'time spent.'
one approach I've considered is using user feedback to steer different activation vectors towards user-preferred responses. that is much closer to the traditional ML paradigm - user feedback updates a signal, which is used at inference time to alter the output of a frozen model. this certainly feels doable (and honestly kinda fun) but challenging without tons of users and scale :)
When the scientists figure out how to enable a remembered collaboration that continually progresses results, then we’ll have something truly remarkable.
I envision an IDE with separate windows for collaborative discussions. One window is the dialog, and another is any resulting code as generated files. The files get versioned with remarks on how and why they've changed in relation to the conversation. The dev can bring any of those files into their project and the IDE will keep a link to which version of the LLM generated file is in place or if the dev has made changes. If the dev makes changes, the file is automatically pushed to the LLM files window and the LLM sees those changes and adds them to the "conversation".
The collaboration continually has "why" answers for every change.
And all of this without anyone digging into models, python, or RAG integrations.
If you build it you'll need to be digging into models, python and various pieces of "RAG" stuff, but your users wouldn't need to know anything about it.
In fact, dynamic evaluation is so old and obvious that the first paper to demonstrate RNNs really worked better than n-grams as LLMs did dynamic evaluation: https://gwern.net/doc/ai/nn/rnn/2010-mikolov.pdf#page=2
Dynamic evaluation offers large performance boosts in Transformers just as well as RNNs.
The fact that it's not standard these days has more to do with cloud providers strongly preferring models to not change weights to simplify deployment and enable greater batching, and preferring to crudely approximate it with ever greater context windows. (Also probably the fact that no one but oldtimers knows it is a thing.)
I think it might come back for local LLMs, where there is only one set of model weights running persistently, users want personalization and the best possible performance, and where ultra-large context windows are unacceptably memory-hungry & high-latency.
The basic idea is that you summarize chunks of the convo over time and create tiny ephemeral datasets that are built for retrieval tasks. You can do this by asking another model to create Q&A pairs for you about the summarized convo context. Each sample in the dataset is an instruction-tuned format with convo context plus the Q&A pair.
The training piece is straightforward but is really where the hair is. Its simple to, in the soft prompt use case, to train a small tensor on that data and just concatenate it to future inputs. But you presumably will have some loss cutoff, and in my experience you very frequently don't hit that loss cutoff in a meaningfully short period of time. Even when you do, the recall/retrieval may still not work as expected (though I was surprised how often it does).
The biggest issue is obviously recall performance given that the weights remain frozen, but also the latency introduced in the real-time training can end up being a major bottleneck, even when you lazily do this work in the background.
If LLMs are actually here to stay, then we need to keep shrinking the inferencing costs to the point that everyone can run it on the client, and avoid all of these problems at once...
It stunned me recently by remembering I had an argon tank, something I had only mentioned in a thread days earlier.
I asked it to recall which conversations these were, for sharing on HN. That conversation was comically stupid (REALLY bad, like "Perplexity" bad) and went nowhere.
Also are we that far from fine tuning in real time? Maybe next year?
I think while the current model might be what he says the model in 2025 could be vastly different.
Plenty of AI vendors will swear that they don't train on input to their models. People straight up don't believe them.
[1] https://www.reuters.com/markets/deals/openai-strikes-deal-br...
1. Most of the data might be junk, and so it's just not worth it
2. They want paying customers, paying customers don't want this, so they don't offer it for paying customers
3. It's really, really hard to annotate data effectively at that scale. The reason why Reddit data is being bought is because it's effectively pre-annotated. Most users don't follow up with chat responses to give feedback. But people absolutely provide answers and upvotes on forums.
As a result, no one trusts the popcorn button, even if it works.
Of course, there’s no real way to know if it’s respected or not.
>All ChatGPT Free users can now use browse, vision, data analysis, file uploads, and GPTs
Both sessions worked toward those implied goals. The one where I asked what was wrong indicated long-winded sentences, redundancy, lack of clear structure, and focus issues. Sorry, Simon!
The one where I asked what was good about the post indicated clarity around misconceptions, good explanations of training, consistent theme, clear sections, logical progression, and focused arguments. I tend to agree with this one more than the other, but clearly the two are conflicted.
So, what's my point?
My point is that AI/LLMs interactions are driven by the human's intent during the interaction. It is literally our own history and training that brings the stateless function responses to where they arrive. Yes, facts and whatnot matter in the training data, and the history, if there is one, but the output is very much dependent on the user.
If we think about this more, it's clear that training data does matter, but likely doesn't matter as much as we think it does. It's probably just as equally important to consider the historic data, and the data coming in from RAG/search processes, as making as big of an impact on the output.
These models have been RLHFed to just go along and agree with whatever the user asks for, even if the user doesn't realize what's going on. This is a bit of a cheap trick- of course users will find something more intelligent if it agrees with them. When models start arguing back you get stuff like early Bing Chat going totally deranged.
Who doesn't already know this?
Seems that's not the case and some think the model could br trained instantly by their input instead of much later when a good amount of new training data is collected.
"Prompts" use symbols within the conversation that's gone before, along with hidden prompting by the organisation that controls access to an LLM, as a part of the prompt. So when you ask, 'do you remember what I said about donuts' the LLM can answer -- it doesn't remember, but that's [an obscured] part of the current prompt issued to the LLM.
It's not too surprising users are confused when purposeful deception is part of standard practice.
I can ask ChatGPT what my favourite topics are and it gives an answer ...
Things usually aren't black and white, unless of course they are bits...