LLMs and the Harry Potter problem
pyqai.com
pyqai.com
Perhaps my biggest gripe: if you're going to lure readers in with an interesting name like the "Harry Potter problem", it better be either technically interesting or entertaining.
Not everything needs to be entertaining to be useful.
LLMs are bad at counting no matter what size of context is provided. If you're going to formulate a thought experiment to illustrate how an LLM stops paying attention well before the context limit, it should be an example that LLMs are known to be good at in smaller context sizes. Otherwise you might be entertaining but you're also misleading.
My hope from this article is to help non-AI experts figure out when they need to design around a flaw versus believe what's marketed.
You're putting a lot of weight into counting. I don't know anyone who wants to use a LLM after hearing "good at math" for counting of all things. Algebra, Calculus, Statistics, hell I used Claude 3 for Special Relativity. Those are the things people will care about when you say math, not counting.
Look, just test your use case and report that lol.
Your article would have been very helpful if you'd simply did that but you didn't so it's not.
https://www.lesswrong.com/posts/qy5dF7bQcFjSKaW58/bad-at-ari...
Edit: Going from counting numbers to finding relevant information at different pages given a task isn't a valid analogy.
Agents are mentioned but not multi-agent architectures [0], where you could have an agent responsible for insurance policies, legal definitions, and/or a bot that is responsible for the big picture in question. They would go back and forth, being expert at their field (or task) and come to a conclusion after some iterations of API calls.
Missed opportunity.
[0] https://microsoft.github.io/autogen/docs/Use-Cases/agent_cha...
The agentic approach falls apart because again, a self-querying mechanism or a multi-agent framework still needs to know where in the document to look for each subset of information. That's why I argue that you need an ontology. And at that point, agents are moot. A small 7b model with a simple prompt suffices, without any of the unreliability of agents. I suggest trying agents on an actually serious document, the problems are pretty evident. That said, I do hope that they get there one day because it will be cool.
Maybe the current state of the art LLM can't solve the kind of high value long context problems you have in mind but what I can tell you though is that you won't find that out by asking it to count.
Write me a short Python program that will count how many times a word occurs in a given text.
Sure, I can help you with that! Here's a short Python program that counts how many times a word occurs in a given text:
```python
def count_words(text, word):
words = > text.lower().split()
count = 0
for w in words:
if w == word.lower():
count += 1
return count
text = "Snapchat is a fun way to share moments with friends!"
word = "Snapchat"
result = count_words(text, word)
print(f"The word '{word}' occurs {result} times in the text.")
```
You can replace the `text` variable with your own text and the `word` variable with the word you want to count. Let me know if you have any questions!
Seems to me that if you give an LLM an environment to run said program it would be able to automatically do this with the correct prompt. This doesn’t solve the insurance policy problem at all but the solution to these problems is different in my opinion.But at the end of the day an LLM is right 80% of the time while being 100% confident 100% of the time that it gets the right answer. You can increase that 80% but I don’t see how the current breed of LLMs can learn to self doubt enough to keep trying to understand better.
Why doesn't the AI itself decide that writing a Python script is a good way to approach the problem?
LLMs (by themselves) cannot reliably count. If you expect them to, then you're falling into the common trap of extrapolating a metacognition layer where none exists.
So you tell me: if a regular developer reads the above, how can they surmise that the model which can do higher-order math can't count?
LLMs do not count. When you ask an LLM to count (which is something it does not do), in the end no counting has happened. No shit, Sherlock.
Claims of context-window sizes increases do imply that information can be extracted from such windows. So this is not trivial.
The counting example is a simple case, the "how much is this insurance policy going to cover this damage" is also a calculation.
An LLM doesn't even answer a question, either: it continues a prompt. That continuation looks like an answer to you and me; but to the LLM, it contains nothing more than the most likely tokens.
I keep hearing this story that an LLM is capable of objectivity, and that it's just relatively bad at it. The LLM literally never does objectivity. It can't be bad at a thing it never does.
Every time someone calls this sort of interaction a "limitation", they are only obfuscating the narrative with a useless anthropomorphization. The LLM is not a person. It is not a mind. It does not perform objective thought. It is a statistical model that provides the most likely text. No more, no less.
"The Harry Potter Problem" has the feel of a strawman. LLMs are not universal problem solvers. You still have to break down tasks or give it a framework for working things through. If you ask an LLM to produce a code snippet for word counting, it will do great. Maybe that isn't as sexy, but what are you really trying to achieve?
> Harry Potter is an innocent example, but this problem is far more costly when it comes to higher value use-cases. For example, we analyze insurance policies. They’re 70-120 pages long, very dense and expect the reader to create logical links between information spread across pages (say, a sentence each on pages 5 and 95). So, answering a question like “what is my fire damage coverage?” means you have to read: Page 2 (the premium), Page 3 (the deductible and limit), Page 78 (the fire damage exclusions), Page 94 (the legal definition of “fire damage”).
It's not at all obvious how you could write code to do that for you. Solving the "Harry Potter Problem" as stated seems like a natural prerequisite for doing this much more high stakes (and harder to benchmark) task, even if there are "better" ways of solving the Harry Potter problem.
LLMs can't count well. This is in large part a tokenization issue. Doesn't mean they couldn't answer all those kind of questions. Maybe the current state of the art can't. But you won't find out by asking it to count.
Not really. The "Harry Potter Problem" as formulated is asking an LLM to solve a problem that they are architecturally unsuited for. They do poorly at counting and similar algorithms tasks no matter the size of the context provided. The correct approach to allowing an AI agent to solve a problem like this one would be (as OP indicates) to have it recognize that this is an algorithmic challenge that it needs to write code to solve, then have it write the code and execute it.
Asking specific questions about your insurance policy is a qualitatively different type of problem that algorithms are bad at, but it's the kind of problem that LLMs are already very good at in smaller context windows. Making progress on that type of problem requires only extending a model's capabilities to use the context, not simultaneously building out a framework for solving algorithmic problems.
So if anything it's the reverse: solving the insurance problem would be a prerequisite to solving the Harry Potter Problem.
The structured output element is really important too - subject for another post though!
So surely it’s not about LLMs being able to do that, but being smart enough to understand “hey this is something I could write a Python script to do” and be able to write the script, feed the Harry Potter chapter into it, run the script and parse the results (in the same way a human would do)?
If you solve the problem that way you don't even really need the Harry Potter chapter in the context at all, you could put it as an external document that the agent executes code against. This makes it qualitatively a different problem than the insurance policy questions that the article moves on to.
What is the relevance in what a human would do? This is not a human and does not work like a human.
I would expect any piece of software that allows me to input a text and ask it to count occurences of words to do so accurately.
This explanation/excuse doesn't hold water.
The large language model's context window absolutely is ephemeral. By the time inference is begun all you have is a giant vector that represents the context to date. This means that the model itself does not have the text available to look at, it only has the encoded "memory" of that text.
OP is simply saying that the underlying model is unsuitable for solving problems like this directly, so it makes a bad example for how models don't use their context effectively. A production grade AI agent should be able to solve problems like this, but it will likely do that through external scaffolding, not through improvements to the model itself, whereas improvements to the context window will probably need to occur at the model level.
It is fairly well-established that context windows are a general issue among LLMs due to SOTA context windows still being somewhere greater than linear. It's also fairly well-established that LLMs aren't necessarily good at things they aren't trained at.
If you are unwilling or unable to throw enough hardware to overcome the context window problem you'll need to reduce the context. If you're unwilling or unable to train the LLM to task you'll have to restructure information such that the task is more tractable.
I'm glad to see that given their constraints they chose a sensible solution for the business, but overall this really seems like a series of known limitations being called out and doesn't feel like it's a good look coming from a company that touts leveraging AI for pulling information from documents and integrating with existing systems...
This is meant to be for the developer who doesn't fit the above profile and thinks a model that has a million token context window and "can handle complex analysis, longer tasks with multiple steps, and higher-order math and coding tasks" (direct quote from Anthropic's website), actually can do those things.
Having read some of your other comments it appears that part of the issue is that you were marketed a 1 million token context window and research has shown that's not quite the case. That said, the article doesn't do a good job of painting that picture - it is alluded to with "all fail at this task despite having big context windows" but I think it's worth being crystal clear here that the marketing says 1m and that is disingenuous in your experience and backed by research findings.
The reason long context windows are advertised is because there are an awful lot of people out there trying to make money replacing customer service agents with LLM-powered chatbots. In order to power them naively (which is the only way sales people who have never built software before know how), you need to feed them a context window full of all your industry/product specific knowledge and then hope that the LLM answers the same way.
But, they don't, and they can't, so you have to spend a lot of time trying to figure out how to tie the LLM down so it responds exactly the way you want, which sort of ruins all the hype and mystery for common people around LLMs. It's been sold as a miracle that can replace people, but the truth is, well, not that, which really hampers the sales process. I think we're seeing this with Elon Musk desperately trying to push people into "FSD". People aren't impressed by AIs that aren't as good as people, if not better than people, at doing whatever task they are supposed to do.
You can accomplish this with RAG.
Your overall point is taken though, the LLM itself is not enough, fine-tuning is not always feasible, and I think no matter how good an AI persona gets at, say, teaching yoga - for some yoga students it will never replace an in-person instructor.
However for a game NPC, online agent, Discord bot, etc. not to mention research, translation, tutorials, summarizing, etc. there is a lot of present day utility for LLMs.
Interesting idea with ragdoll, but I'd hate to try to compete with Tavern/Pygmalion cards + Lorebooks, seems like there is a critical mass there already for RP chatbots.
There is a chat mode where the UI looks like a chat (communication is key with these guys), but there are also views like Picture mode where you paste images to create concept art or upload a still to generate a movie from - so it's more like a creative studio with different views (one of which is chat).
The best part of it all is you're like the boss or conductor, directing a staff of 1 or 2 or hundreds of AI personas all with specific knowledge and abilities to create whatever you need.
It doesn't matter if the context window is large or small, the Harry Potter Problem as formulated is going to be just as hard because it's not a problem with false advertising in context window sizes, it's a problem inherent to the computing paradigm.
A version of the Harry Potter Problem that was formulated around a model's ability to recall specific scenes of a novel would be much more useful as an illustration of the limitations of the supposedly-large context windows.
And if I can't trust a so-called SOTA model to partially answer - say, recall each mention of the word "wizard" instead of just giving me the wrong answer - then why should I trust it to list out specific scenes? That's even harder to benchmark.
How soon until the hype fades and we enjoy our next AI winter?
>counting
Pick one.
This is very similar to the "precision" misconception regarding floating point numbers.
The answer isn't wrong, it's just imprecise.
Hallucinations are a misnomer.
You are trying to get exact integer<>word accuracy from an architecture that is innately probabilistic, and where atomically it clashes; words get tokenized, so arithmetic is difficult at a microscale - the carry bit likely won't make it to the (needed transformer) context to work, since usually, most numbers don't overflow on average when summed.
It can, however, output a small program - with high confidence - that it can self-evaluate for functional proximity, then use that to help arrive at an answer.
This is a proto-Mixture of Experts model, achieved by another hyper-visor or guard dog LLM.
The point here is: the justifications from AI engineers for why counting vs math aren't the same task, while valid, are irrelevant because marketing never brings up the limitation in the first place. So any logical person who doesn't know a lot about AI will arrive at a logical, albeit practically incorrect conclusion.
But that's not what they said; to be fair. They said it can do complex math - not simple math, repeatedly, many times, by one inference.
The architecture just clashes against the intent too much to arrive at a useful/acceptable answer.
Had you crafted a larger prompt that recursively divides the context into n amount of separation buckets, then sum them (inverted binary tree wise), you'd likely have better luck with the carry bits tallying correctly.