What We Learned from a Year of Building with LLMs
oreilly.com
oreilly.com
1) you’re sampling a distribution; if you only sample once, your sample is not representative of the distribution.
For evaluating prompts and running in production; your hallucination rate is inversely proportional to the number of times you sample.
Sample many times and vote is a highly effective (but slow) strategy.
There is almost zero value in evaluating a prompt by only running it once.
2) Sequences are generated in order.
Asking an LLM to make a decision and justify its decision in that order is literally meaningless.
Once the “decision” tokens are generated; the justification does not influence them. It’s not like they happen “all at once” there is a specific sequence to generating output where the later output cannot magically influence the output which has already been generated.
This is true for sequential outputs from an LLM (obviously), but it is also true inside single outputs. The sequence of tokens in the output is a sequence.
If you’re generating structured output (eg json, xml) which is not sequenced, and your output is something like {decision: …, reason:…} it literally does nothing.
…but, it is valuable to “show the working out” when, as above, you then evaluate multiple solutions to a single request and pick the best one(s).
To the user
But these tools are marketed as if you do only need to run them once to get a good result; The companies behind them would really want you to stop hammering the button that deletes their money.
As an aside:
> For evaluating prompts and running in production; your hallucination rate is inversely proportional to the number of times you sample.
This isn't really true, and requires you to fuzz the prompt itself for best effect. Making the "spam the LLM with requests" problem much worse.
Is this true if you are using RAG too?
This isn't correct.
You're just sampling a different distribution.
You can adjust the shape of the distribution with your prompt; certainly... and if you make a good prompt, perhaps, you can narrow the 'solution space' that you sample into.
...but, you're still sampling randomly into a distribution, and the N'th token relies on the (N-1)'th token as an input; that means that a random deviance to a bad solution is incrementally responsible for a bad solution, regardless of your prompt.
...
Consider the prompt "Your name is Pete. What is your name?"
Seems like a fairly narrow distribution right?
However, there's a small chance that the first generated token is 'D'; it's small, but non-zero. That means it happens from time to time. The higher the temperature, the higher the randomization of the output tokens.
How do you imagine that completion runs when it happens? Doug? Dane? Danial? Dave? Don't know? I tell you what it is not; it's not Pete.
That's the issue here; when you sample, the solution space is wide, and any single sample has a P chance of being a stupid hallucination.
When you sample multiple times, the chance of that hallucination is P * P * P * P, etc. by the number of time you sample.
You can therefore control your error rate this way, because, you can calculate the chance of failure as P^N.
Yes, obviously, if your P(good answer) < P(bad answer) it has the opposite effect.
...but no, sampling once with a single prompt does not save you from this prompt no matter what or how good your prompt is.
Furthermore, when you evaluating prompts, only sampling once means you have no way of knowing if it was a good prompt or not. While, if you sample say, 10 times, you can see that obviously, from the outputs (eg. Pete, Pete, Pete, Pete, Potato, Pete, Pete <--- ) what the prompt is doing.
You can measure the error rate in your prompts this way.
If you don't, honestly, you really have no idea if your prompts are any good at all. You're just guessing.
People who run a prompt, tweak it, run it, tweak it, run it, tweak it, etc. are observing random noise, not doing prompt engineering.
Never sample only once.
Edit in response to your wall of text: I have *extensively* tested the results of multi-shot prompting vs repeated single shot prompting, and the differences between them are not material to the outcome of "averaging" results, or selecting the best result. You can theorize all you want, but the real world would like a word.
It's an easy adjustment to workflows that people often either forget to do or don't realise they should be doing.
Sorry; I'm not trying to criticize; I'm just telling you that's how it works.
OP is saying that you can’t evaluate any prompt by using just one generation using that prompt.
You need to run that prompt several times to approximate any prompt’s performance. That’s just how probability works.
I don’t believe OP is arguing the effectiveness of running 10 different prompts vs a single multiple perspective prompt.
In your example you will get Pete every time.
I will say though, using temperature 0 without understanding it (or worse, testing at temp > 0 and then setting temp to 0 for production, which I literally had to stop someone I know and respect as a developer from doing) and not understanding what top_k and top_n do (but using them anyway) is my #3 for LLM fails.
/shrug
...but, yes as you say, in a trivial case, like binary decision making, a 0 or very low temperature can help with the need to multiple sample; and as you say, when it's deterministic, sampling multiple times doesn't help at all.
Beam search[1] has long been a great way to sample from language models, even before transformers. Essentially you keep track of the top N most promising threads and sample randomly from those.
OpenAI doesn't offer beam search yet, just temperature and top_k, but I hope they add support for it because it's far more efficient than just starting over each time.
1) Hallucination rate is not inversely proportional to number of samples, unless you assume statistical independence. As you’re sampling from the same generative process each time, any inherent bias of the LLM could affect every sample (eg see golden gate Claude). Naively calculating hallucination rate as P^N is going to be a massive underestimate of the true error rate for many tasks requiring factual accuracy.
2) You’re right that output tokens are generated autoregressively, but you are thinking like a human. Transformer attention layers are permutation invariant. The ordering of output (eg decision first then justification later) is inconsequential, either can be derived from input context and hidden state where there is no causal masking of attention.
With decision before justification, you tend to have a greater risk of the output being a wrong decision followed by convincing BS justifying it.
(edit: Another way you could think of it is, LLMs still can't violate causality. Attention heads' ability to look in both directions with respect to a particular token's position in the sequence does not enable them to see into the future and observe tokens that don't exist yet.)
This is crazy town banana pants.
Point 1 is also a good callout. I added something on this for llm judge but it’s relevant more broadly.
The LLM is like another user. And it can surprise you just like a user can. All the things you've done over the years to sanitize user input apply to LLM responses.
There is power beyond the conversational aspects of LLMs. Always ask, do you need to pass the actual text back to your user or can you leverage the LLM and constrain what you return?
LLMs are the best tool we've ever had for understanding user intent. They obsolete the hierarchies of decision trees and spaghetti logic we've written for years to classify user input into discrete tasks (realizing this and throwing away so much code has been the joy of the last year of my work).
Being concise is key and these things suck at it.
If you leave a user alone with the LLM, some users will break it. No matter what you do.
I like to think of LLMs as client-side code, at least in terms of their risk-profile.
No data you put into them (whether training or prompt) is reliably hidden from a persistent user, and they can also force it to output what they want.
I really like this analogy! That sums up my experiences with LLMs as well.
(Note: this is only Part 1 of 3 of a series that has already been written and the other 2 parts will be released shortly)
Sure. Ultimately, you want to use KG to increase your ability to do great retrieval.
Why do graphs help with retrieval? Well, don’t overlook the classic pageant example: graphs provide signal about the interconnectivity of the docs.
Also, sometimes the graph itself are a kind of object you want to retrieve over.
If I'm understanding this correctly, the standard way to get structured output seems to be to retry the query until the stochastic language model produces expected output. RAG also seems like a hilariously thin wrapper over traditional search systems, and it still might hallucinate in that tiny distance between the search result and the user. Like we're talking about writing sentences and coaching what amounts to an auto complete system to magically give us something we want. How is this industry getting hundreds of billions of dollars in investment?
Also the error rate is about 5-10% according to this article. That's pretty bad!
FOMO? To me it's the Gold Rush, except that it's not clear if anyone wants that kind of gold at the end :-).
Having 90-95% success rate on something that was previously impossible is acceptable. Without LLMs the success rate would be 0% for the things I'm doing.
It's probable that Google is in the middle of doing this napkin math given all the embarrassing stuff we saw last week. So it's cool that we're closer to solving these really hard problems but whether they're acceptable is a more complicated question than just it used to not be possible. Maybe that math works out in your favor for your application.
These systems compete with humans, not with formatters.
Even if you had to implement checks and balances for AI systems, you'd still come away having spent way less money.
No, that would be very inefficient. At each token generation step, the LLM provides a likelihood for all the defined token based on the past context. The structured output is defined by a grammar, which defines the legal tokens for the next step. You can then take the intersection of both (ignore any token not allowed by the grammar), and then select among the authorized token based on the LLM likelihood for them in the usual way. So it's a direct constraint, and it's efficient.
Also, what's it for? None of these articles point to anything worthwhile that it's useful for.
I found that the simpler the better, when testing lots of different SQL schema formats on https://www.sqlai.ai/. CSV (table name, table column, data type) outperformed both a JSON formatted and SQL schema dump. And not to mention consumed fewer tokens.
If you need the database schema in a consistent format (e.g. CSV) just have LLM extract data and convert whatever the user provides into CSV. It shines at this.
I found that similarly-named columns easily confused GPTs
I am curious to know if the authors tried to build LLMs in languages other than English and what did they learn while doing so?
An excellent post reminding me of the best O'Reilly articles from the past. Looking forward to parts 2 and 3.
I didn’t realize at first you meant to highlight a language barrier — being well structured is a challenge for most in their native tongue!
One controversial point that has led to discussions in my team is this:
> A common anti-pattern/code smell in software is the “God Object,” where we have a single class or function that does everything. The same applies to prompts too.
In theory, a monolithic agent/prompt with infinite context size, a large toolset, and perfect attention would be ideal.
Multi-agent systems will always be less effective and more error-prone than monolithic systems on a given problem because of less context of the overall problem. Individual agents work best when they have entirely different functionalities.
I wrote down my thoughts about agent architectures here: https://www.kadoa.com/blog/ai-agents-hype-vs-reality
If each step of your task requires knowledge of the big picture, then yeah it ought to help to put all your context into a single API call.
But if you can decompose your task into relatively independent subtasks, then it helps to use a custom prompt/custom model for each of those steps. Extraneous context and complexity are just opportunities for the model to make mistakes, and the more you can strip those out, the better. 3 steps with 99% reliability are better than 1 step with 90% reliability.
Of course, it all depends on what you're trying to do.
I'd say single, big API calls are better when:
- Much of the information/substeps are interrelated
- You want immediate output for a user-facing app, without having to wait for intermediate steps
Multiple, sequenced API calls are better when:
- You can decompose the task into smaller steps, each of which do not require full context
- There's a tree or graph of steps, and you want to prune irrelevant branches as you proceed from the root
- You want to have some 100% reliabile logic live outside of the LLM in parsing/routing code
- You want to customize the prompts based on results from previous steps
Write the initial prompt, and write some tests to validate the output of the prompt. Then as the prompt grows you can observe the decline in performance at the initial task. Then you can decide if the new larger prompt is worth the decline in performance at it’s initial task.
It might not work for all tasks, but a good candidate would be - write SQL queries from natural language.
should be fun!
Some notes from my own experience on LLMs for NLP problems:
1) The output schema is usually more impactful than the text part of a prompt.
a) Field order matters a lot. At inference, the earlier tokens generated influence the next tokens.
b) Just have the CoT as a field in the schema too.
c) PotentialField and ActualField allow the LLM to create some broad options and then select the best. This mitigates the fact that they can't backtrack a bit. If you have human evaluation in your process, this also makes it easier for them to correct mistakes.
`'PotentialThemes': ['Surreal Worlds', 'Alternate History', 'Post-Apocalyptic'], 'FinalThemes': ['Surreal Worlds']`
d) Most well definined problems should be possible zero-shot on a frontier model. Before rushing off to add examples really check that you're solving the correct problem in the most ideal way.
2) Defining the schema as typescript types is flexible and reliable and takes up minimal tokens. The output JSON structure is pretty much always correct (as long as the it fits in the context window) the only issue is that the language model can pick values outside the schema but that's easy to validate in post.
3) "Evaluating LLMs can be a minefield." yeah it's a pain in the ass.
4) Adding too many examples increases the token costs per item a lot. I've found that it's possible to process several items in one prompt and, despite it being seemingly silly and inefficient, it works reliably and cheaply.
5) Example selection is not trivial and can cause very subtle errors.
6) Structuring your inputs with XML is very good. Even if you're trying to get JSON output, XML input seems to work better. (Haven't extensively tested this because eval is hard).
Love the idea of adding CoT as a field in the expected structured output as it also makes it easier from a UX perspective to show/hide internal vs external outputs.
> Structuring your inputs with XML is very good. Even if you're trying to get JSON output, XML input seems to work better. (Haven't extensively tested this because eval is hard).
Would be neat to see LLM-specific adapters that can be used to swap out different formats within the prompt.
In places that grew up learning UK English we use delve not that dissimilar to ChatGPT.
it's originally "Ready to ~~delve~~ dive in?" but something got lost in translation
Is it the thing itself, or is it the thing that enables us?