What we've learned from a year of building with LLMs
eugeneyan.com
eugeneyan.com
Part 1: https://www.oreilly.com/radar/what-we-learned-from-a-year-of... Part 2: https://www.oreilly.com/radar/what-we-learned-from-a-year-of...
We were working on this webpage to collect the entire three part article in one place (the third part isn't published yet). We didn't expect anyone to notice the site! Either way, part 3 should be out in a week or so.
I love this new era of computing we're in where rumors, second-guessing and something akin to voodoo have entered into working with LLMs.
That’s why LLMs are good at translating and spellchecking. We’ve been describing the same world and almost all texts respect grammar. That’s the first things that surface. But you can extract the same rules in other way and create a program that does it without the waste of computing power.
If we describe computing as solving problems, then it’s not computing because if your solution was not part of the training data, you won’t solve anything. If we describe computing as symbol manipulation, then it’s not doing a good job because the rules changes with every model and they are probabilistic. No way to get a reliable answer. It’s divination without the divine (no hint from an omniscient entity).
Imagine if physics literature was filled with stuff about psychology and how that would drive physicists nuts. That's how I feel right now ;)
If I can do node-red or a function chain for prompts and outputs, that would be sweet.
“You are in charge of game prep and must work with an LLM over many prompts to…”
Instead of trying to do everything into a single chat or chain, add steps to ask the LLM to break down the next tasks, with context, and store that into SQLite or something. Then start new chats/chains on each of those tasks.
Then just loop them back into LLM.
I find that long chats or chains just confuse most models and we start seeing gibberish.
Right now I'm favoring something like:
"We're going to do task {task}. The current situation and context is {context}.
Break down what individual steps we need to perform to achieve {goal} and output these steps with their necessary context as {standard_task_json}. If the output is already enough to satisfy {goal}, just output the result as text."
I find that leaving everything to LLM in a sequence is not as effective as using LLM to break things down and having a DB and code logic to support the development of more complex outcomes.
Also mentioning what to "forget" or not focus on anymore seems to remove some noise from the responses if they are large.
https://dho.stanford.edu/wp-content/uploads/Legal_RAG_Halluc...
So would strongly disagree that LLMs have become “good enough” for real-world applications" based on what was promised.
Disclosure: author on [1]
It's the yes we hallucinate but don't worry because we provide the sources for users to check.
Even though everyone knows that users will never check unless the hallucination is egregious.
It's such a disingenuous way of handling this.
I can't speak for "what was promised" by anyone, but LLMs have been good enough to live in production as a core feature in my product since early last year, and have only gotten better.
I would love to see this article also expand to touch upon things like : - data management - (tooling, frameworks, open vs closed data management, labelling & annotations) - inference as a pipeline - frameworks for breaking down model inference into smaller tasks & combining outputs (do DAG's have a role to play here?) - prompts - areas like caching, management, versioning, evaluations - model observability - tokens, costs, latency, drift? - evals for multimodality - how do we tackle evals here which in turn can go into loops e.g. quality of audio, speech or visual outputs
Me personally, I only used LLM for one "serious" application: I used GPT-3.5Turbo for transforming unstructured text into JSON; it was basically just ad-hoc Node.js script that called API (prompt was few examples of input-output pairs), and then it did some checks (these checks usually failed only because GPT also corrected misspellings). It would take me weeks to do it manually, but with the help of GPT it was few hours (writing of the script + I made a lot of misspellings so the script stopped a lot). But I cannot imagine anything more complex.
Since you seem to have not noticed my comment above, here's another example of a project that implements many of these techniques. Me and many others have used this to transcribe hour long videos into a well organized "docs site" that makes the content easy to read.
Example: https://matadoc.vercel.app/
This was completely auto-generated in a few minutes. The author of the library reviewed it and said that it's nearly 100% correct and people in the company where it was built rely on these docs.
Tell me how long it would take you to write these docs. I'm really confused where your dismissive mentality is coming from in the face of what I think is overwhelming evidence to the contrary. I'm happy to provide example after example after example. I'm sorry, but you are utterly, completely wrong in your conclusions.
What you're describing is more about reasoning abilities - that's not really what the article was about or the problems the techniques are for. The techniques in article are more for stuff like Q&A, classification, summarization, etc.
Here's a more dramatic example: https://www.grey-wing.com/
This company provides deeply integrated LLM-powered software for operating freight ships.
There are a lot of people who are doing this and achieving very good results.
Sorry, if it's not working for you, it doesn't mean that it doesn't work.
Weird: I just refreshed the page and it now redirects to a different domain (than the originally-submitted URL) and has a date of June 8, 2023. It still cites articles and blog posts from 2024, though.
Here is a summary of all points:
1. Focus on Prompting Techniques:
1.1. Start with n-shot prompts to provide examples demonstrating tasks.
1.2. Use Chain-of-Thought (CoT) prompting for complex tasks, making instructions specific.
1.3. Incorporate relevant resources via Retrieval Augmented Generation (RAG).
2. Structure Inputs and Outputs: 2.1. Format inputs using serialization methods like XML, JSON, or Markdown.
2.2. Ensure outputs are structured to integrate seamlessly with downstream systems.
3. Simplify Prompts: 3.1. Break down complex prompts into smaller, focused ones.
3.2. Iterate and evaluate each prompt individually for better performance.
4. Optimize Context Tokens: 4.1. Minimize redundant or irrelevant context in prompts.
4.2. Structure the context clearly to emphasize relationships between parts.
5. Leverage Information Retrieval/RAG: 5.1. Use RAG to provide the LLM with knowledge to improve output.
5.2. Ensure retrieved documents are relevant, dense, and detailed.
5.3. Utilize hybrid search methods combining keyword and embedding-based retrieval.
6. Workflow Optimization: 6.1. Decompose tasks into multi-step workflows for better accuracy.
6.2. Prioritize deterministic execution for reliability and predictability.
6.3. Use caching to save costs and reduce latency.
7. Evaluation and Monitoring: 7.1. Create assertion-based unit tests using real input/output samples.
7.2. Use LLM-as-Judge for pairwise comparisons to evaluate outputs.
7.3. Regularly review LLM inputs and outputs for new patterns or issues.
8. Address Hallucinations and Guardrails: 8.1. Combine prompt engineering with factual inconsistency guardrails.
8.2. Use content moderation APIs and PII detection packages to filter outputs.
9. Operational Practices: 9.1. Regularly check for development-prod data skew.
9.2. Ensure data logging and review input/output samples daily.
9.3. Pin specific model versions to maintain consistency and avoid unexpected changes.
10. Team and Roles: 10.1. Educate and empower all team members to use AI technology.
10.2. Include designers early in the process to improve user experience and reframe user needs.
10.3. Ensure the right progression of roles and hire based on the specific phase of the project.
11. Risk Management: 11.1. Calibrate risk tolerance based on the use case and audience.
11.2. Focus on internal applications first to manage risk and gain confidence before expanding to customer-facing use cases.- "Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?" https://arxiv.org/abs//2405.05904
- "Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs" https://arxiv.org/abs/2312.05934
But "knowledge injection" is still pretty narrow to me. Here's an example of a very simple but extremely valuable usecase - taking a model that was trained on language+code and finetuning it on a text-to-DSL task, where the DSL is a custom one you created (and thus isn't in the training data). I would consider that close to infeasible if your only tool is a RAG hammer, but it's a very powerful way to leverage LLMs.
"Teaching" the LLM an entirely new language (like a DSL) might actually need fine-tuning, but you can probably build a pretty decent first-cut of your system with n-shot prompts, then fine-tune to get the accuracy higher.
As with other situations that want a custom DSL, our syntax has its own quirks and details, but is similar enough to e.g. Mermaid that we are able to produce valid syntax pretty easily.
What we've found harder is controlling for edge cases about how to build proper diagrams.
For more context: https://www.eraser.io/decision-node/on-building-with-ai
As for lora - in the context of my comment, that's just splitting hairs IMO. It falls in the category of finetuning for me, although I understand why you might disagree. But it's not like the article mentions lora either, nor am I aware of people doing lora without GPUs which the article is against (No GPUs before PMF)
A concrete example is something like a tool for querying an internal company wiki and the query "tell me about the Backend Team's approach to sprint planning". Normal retrieval approaches will pull information directly related to that query. But what if there is no information about the backend team's practices? As a human, you would do multi-hop/horizontal information extraction - you would retrieve information about who makes up the backend team, you would then retrieve information about them and their backgrounds/practices. You might might have a hypothesis that people carry over their practices from previous experiences, so you look at the previous teams and their practices. Then you would have the context necessary to give a good answer. I don't know of many people implementing RAG like that. And what I described is 100% possible for AI to do today.
Techniques that would get around this like iterative retrieval or retrieval-as-a-tool don't seem popular.
Its not a technical issue, its a practicality issue imo.
Sometimes RAG is enough. Sometimes fine tuning on top of RAG is better. It depends on the use case. I can't think of any examples where you would want to fine tune and not use rag as well.
Sometimes you fine tune a small model so it performs close to a larger varient on that specific narrow task and you improve inference performance by using a smaller model.
Note that their guidance here is quite practical:
> If prompting gets you 90% of the way there, then fine-tuning may not be worth the investment.
Expensive? Sure, all of AI is crazy expensive. Unfeasible? No
All things considered, in relative terms, as much as I think fine-tuning would be nice, it will remain significantly more expensive than just making RAG or search calls. I say this while being a fan of fine-tuning.
Going to have to disagree with you on that one. A modern 8B model that has been trained on enough tokens is ridiculously powerful.
Don't get me wrong. I think an 70B or larger model would be worth fine-tuning, especially if it can be grown further with more layers.
Any evidence of that that I can look at? This doesn't match what I've seen nor have I heard this from the world-class researchers I have worked with. Would be interested to learn more.
- Most of our accuracy ROI is from agentic loops over top models, and dynamic RAG example injection goes far here that the relative lift of adding fine-tuning isn't worth the many costs
- A lot of fine-tuning is for OSS models that do worse than agentic loops over the proprietary GPT4/Opus3
- For distribution, it's a lot easier to deploy for pluggable top APIs without requiring fine-tuning, e.g., "connect to your gpt4/opus3 + for dumber-but-bigger tasks, groq"
- The resources we could put into fine-tuning are better spent on RAG, agentic loops, prompts/evals, etc
We do use tuned smaller dumber models, such as part of a coarse relevancy filter in a firehose pipeline... but these are outliers. Likewise, we expect to be using them more... but again, for rarer cases and only after we've exhausted other stuff. I'm guessing as we do more fine-tuning, it'll be more on embeddings than LLMs, at least until OSS models get a lot better.
The subtle bit is just doesn't have to be for LLMs, as these are typically part of a system-of-models. E.g., we <3 RAG, and GNNs for improving your KG is fascinating. Likewise, dspy's explorations in optimizing prompts, vs LLMs, is very cool.
Oh man I am so torn between this being a fantastic idea and this being "building a better slide-rule in the age of the computer".
dspy is definitely a project I want to dig into more
Once RAG projects become important and good answers matter - we work with governments, manufacturers, banks, cyber teams, etc - working through data quality, data representation, & retrieval quality helps
Note that we didn't start here: We began with naive RAG, then relevancy filtering, then agentic & neurosymbolic querying, then dynamic example prompt injection, and now are getting into cleaning up the database/kg itself
For folks doing investigative/analytics projects in this space, happy to chat about what we are doing w Louie.AI. These are more implementation details we don't normally write about.
Identifying misinfo - Ranking & summarization based on internet data should be a lot more careful, and sometimes the controversy is the interesting part
For both, GNNs are generally SOTA
A year+ later, the most interesting kernel of insight to us from dspy is autotuning a single prompt: it's an optimizeable model just like any other. As soon as you have an eval framework in place for your prompts, having something like dspy tune your prompts on a per-LLM basis would be very cool. I'm not sure where they are on that, it seems against the grain for their focus. We're only now reaching the point where we would see ROI on that kind of thing, it took a long time to get here.
We do run an agentic framework, so doing cross-prompt autotuning would be neat too -- especially for how the orchestrator (ex: CoT) composes with individual agents. We call this the "composition problem" and it's frustrating. However, again, dspy and friends do "too much", by trying to also be the agent framework & runtime, while we just want the autotuner.
Again... I'm truly happy and supportive that academics are exploring a wild side of the design space. Just, as we are in the 'we ship code people rely on' side of the universe, it's hard to find scenarios where its potential benefits outweigh its costs.
The article has a section called "When to finetune", along with links to separate pages describing how to do so. They absolutely don't say that "fine-tuning isn't even a consideration". Instead, they describe the situations in which fine-tuning is likely to be helpful.
I’d you’re interested in using one of the LLM-applications I have in prod, check out https://hex.tech/product/magic-ai/ It has a free limit every month to give it a try and see how you like it. If you have feedback after using it, we’re always very interested to hear from users.
As far as fine-tuning in particular, our consensus is that there are easier options first. I personally have fine-tuned gpt models since 2022; here’s a silly post I wrote about it on gpt 2: https://wandb.ai/wandb/fc-bot/reports/Accelerating-ML-Conten...
I went back while writing this comment and realized it might be showing me a diff (better use of color would have helped, I have been trained by github). But I was at a loss for what to do with that. I just now figured out the Keep button exists and it accepted the diff and now it sort of makes sense, but the SQL still doesn't return any results.
My honest feedback is that there is way too much stuff I don't understand on the screen and it makes me confused and a little stressed. Ease me into it please, I'm dumb. There seems to be cells that are linked together and cells that aren't(? separated by purplish background) and I don't understand it. I am a jupyter user and I feel like this should be intuitive to me, but it isn't. I am not a designer, but I suspect the structural markings like cell boundaries are too faint compared to the content of the cells and/or the exterior of a cell having the same color as the interior is making it hard for me. I feel lost in a sea of white.
But the core issue is that, excluding the prompt I copy-pasted word for word which worked like a charm, I am 0 out of 4 on actually leveraging AI to solve the problems I asked of Magic. I like the concept of natural language BI (I worked on in the early days when Alexa came out) so I probably gave it more chances than I would have for a different product.
For me, it doesn't fit my criteria for good problems to solve with AI in 2024 - the conversational interface and binary right/wrong nature of querying/presenting data accurately make the cost of failure too high, which is a death sentence for AI products IMO (compare to proactive, non-blocking products like copilot or shades-of-wrong problems like image generation or conversations with imaginary characters). But text-to-SQL and data presentation make sense as AI capabilities in 2024 so I can see why that could be a good product to pursue. If it worked, I would definitely use it.
> June 8, 2024
Is this an article from the future?
Best guess is that's the anticipated publishing date of the full three parts on the official O'Reilly site.
Thanks a lot for this!
The problem I see is, who can an "application" be anything but a little window onto the base abilities of ChatGPT and so effectively offers nothing more to an end-user. The final result still have to be checked and regular end-users have to do their own prompt.
Edit: Also, I should also say that anyone who's designing LLM apps that, rather than being end-user tools, are effectively gate keepers to getting action or "a human" from a company deserves a big "f* you" 'cause that approach is evil.
-Regex expressions: ChatGPT is the best multi-million regex parser to date.
-Grammar and semantic check: It's a very good revision tool, helped me a lot of times, specially when writing in non-native languages.
-Artwork inspiration: Not only for visual inspiration, in the case of image generators, but descriptive as well. The verbosity of some LLMs can help describe things in more detail than a person would.
-General coding: While your mileage may vary on that one, it has helped me a lot at work building stuff on languages i'm not very familiar with. Just snippets, nothing big.
- Generate targeted LLM micro summaries of every record (ticket, call, etc.) continually
- Use layers of regex, semantic embeddings, and scoring enrichments to identify report rows (pivots on aggregates) worth attention, running on a schedule
- Proactively explain each report row by identifying what’s unusual about it and LLM summarizing a subset of the microsummaries.
- Push the result to webhook
Lack of JSON schema restriction is a significant barrier to entry on hooking LLMs up to a multi step process.
Another is preventing LLMs from adding intro or conclusion text.
How are you struggling with this, let alone as a significant barrier? JSON adherence with a well thought out schema hasn't been a worry between improved model performance and various grammar based constraint systems in a while.
> Another is preventing LLMs from adding intro or conclusion text.
Also trivial to work around by pre-filling and stop tokens, or just extremely basic text parsing.
Also would recommend writing out Stream-Triggered Augmented Generation since the term is so barely used it might as well be made up from the POV of someone trying to understand the comment
You work around it with post-processing and retries. But it’s still a bit brittle given how much stuff happens downstream without supervision.
Out of curiosity- do those orgs not find the loss of generality that comes from custom models to be an issue? e.g. vs using Llama or Mistral or some other open model?
Might not need JSON but whatever format it outputs, it needs to be reliable.
I was parsing a document recently, 10-ish questions for 1 document, would make things expensive.
Might be what’s needed but not ideal.
Seems like it would be universally useful.
If you're using anything less you should have a grammar that enforces exactly what tokens are allowed to be output. Fine Tuning can help too in case you're worried about the effects of constraining the generation, but in my experience it's not really a thing
Might be worth checking out.
(Plug) I shipped a dedicated OpenAI-compatible API for this, jsonmode.com a couple weeks ago and just integrated Groq (they were nice enough to bump up the rate limits) so it's crazy fast. It's a WIP but so far very comparable to JSON output from frontier models, with some bonus features (web crawling etc).
You can check it out over at https://github.com/BoundaryML/baml. Would love to talk if this is something that seems interesting!
This is really interesting, is there any architecture documentation/articles that you can recommend?
https://www.linkedin.com/pulse/ai-2024-more-answers-fewer-qu...
I often see these messages from the community doubting the reality, but LLMs are a powerful tool in the tool chest. But I think most companies are not staffed with skilled enough engineers with a creative enough bent to really take advantage of them yet or be willing to fund basic research and from first principles toolchain creation. That’s ok. But it’s foolish to assume this is all hype like crypto was. The parallels are obvious but the foundations are different.
But the facts are that today LLMs are not suitable for use cases that need accurate results. And there is no evidence or research that suggests this is changing anytime soon. Maybe for ever.
There are very strong parallels to crypto in that (a) people are starting with the technology and trying to find problems and (b) there is a cult like atmosphere where non-believers are seen as being anti-progress and anti-technology.
On the crypto stuff yeah I get it - especially if you’re not in the weeds of its use. A lot of people formed opinions from GPT3.5, Gemini, copilot, and other crappy experiences and haven’t kept up with the state of the art. The rate of change in AI is breathtaking and I think hard to comprehend for most people. Also the recent mess of crypto and the fact grifters grift etc also hurts. But people who doubt -are- stuck in the past. That’s not necessarily their fault and it might not even apply to their career or lives in the present and the flaws are enormous as you point out. But it’s such a remarkably powerful new mode of compute that it in combination with all the other powerful modes of compute is changing everything and will continue too, especially if next generation models keep improving as they seem to be likely to.
To me it still looks like a hammer made completely from rubber. You can practice to get some good hits, but it is pretty hard to get something reliable. And a beginner will basically just bounce it around. But it is sold as rescue for beginners.
That sounds like corporate buzzword salad. It doesn't tell much as it stands, not without at least one specific example to ground all those relative statements.
However, I have two that do, which I've discussed in the article. These are two production use cases that I have supported (which again, are explicitly mentioned in the article):
1. https://www.honeycomb.io/blog/introducing-query-assistant
2. https://www.youtube.com/watch?v=B_DMMlDuJB0
Other co-authors have worked on significant bodies of work:
Bryan Bischoff lead the creation of Magic in Hex: https://www.latent.space/p/bryan-bischof
Jason Liu created the most popular OSS libraries for structured data called instructor https://github.com/jxnl/instructor, and works with some of the leading companies in the space like Limitless and Raycast (https://jxnl.co/services/#current-and-past-clients)
Eugene Yan works with LLMs extensively at Amazon and uses that to inform his writing: https://eugeneyan.com/writing/ (However he isn't allowed to share specifics about Amazon)
I believe you might find these worth looking at.
But those do not seem to be real world business cases.
Can you expand a bit more why you think they are? We don't have hours to spend reading, and you say you've been allowed to talk about them.
So can you summarise the business benefits for us, which is what people are asking for, instead of linking to huge articles?
There's a summary for ya! More details in the stuff that they linked if you want to learn. Technical skills do require a significant time investment to learn, and LLM usage is no different.
The first one is a real world product that lives in production that is user facing for a paid product.
The second video goes in depth about how a AI assistant was built for a real estate CRM company, also a paid product.
I don’t understand the assertion that it’s not “real world” or not “business”
Here are additional articles about these
https://help.rechat.com/guides/lucy
https://www.prnewswire.com/news-releases/honeycomb-launches-...
Hello! I'm the owner of the feature in question who experimented with chatgpt last year in the course of building the feature (and working with Hamel to improve it via fine-tuning later).
Even today, it could not work with ChatGPT. To generate valid queries, you need to know which subset of a user's dataset schema is relevant to their query, which makes it equally a retrieval problem as it does a generation problem.
Beyond that, though, the details of "what makes a good query" are quite tricky and subtle. Honeycomb as a querying tool is unique in the market because it lets you arbitrarily group and filter by any column/value in your schema without pre-indexing and without any cost w.r.t. cardinality. And so there are many cases where you can quite literally answer someone's question, but there are multitudes of ways you can be even more helpful, often by introducing a grouping that they didn't directly ask for. For example, "count my errors" is just a COUNT where the error column exists, but if you group by something like the HTTP route, the name of the operation, etc. -- or the name of a child operation and its calling HTTP route for requests -- you end up actually showing people where and how these errors come from. In my experience, the large majority of power users already do this themselves (it's how you use HNY effectively), and the large majority of new users who know little about the tool simply have no idea it's this flexible. Query Assistant helps them with that and they have a pretty good activation rate when they use it.
Unfortunately, ChatGPT and even just good old fashioned RAG is often not up to the task. That's why fine-tuning is so important for this use case.
However, I have two that do, which I've discussed in the article. These are two production use cases that I have supported (which again, are explicitly mentioned in the article):
1. https://www.honeycomb.io/blog/introducing-query-assistant
2. https://www.youtube.com/watch?v=B_DMMlDuJB0
Other co-authors have worked on significant bodies of work:
Bryan Bischoff lead the creation of Magic in Hex: https://www.latent.space/p/bryan-bischof
Jason Liu created the most popular OSS libraries for structured data called instructor https://github.com/jxnl/instructor, and works with some of the leading companies in the space like Limitless and Raycast (https://jxnl.co/services/#current-and-past-clients)
Eugene Yan works with LLMs extensively at Amazon and uses that to inform his writing: https://eugeneyan.com/writing/ (However he isn't allowed to share specifics about Amazon)
I believe you might find these worth looking at.
For example, we focused on the boring and hard task of web data extraction.
Traditional web scraping is labor-intensive, error-prone, and requires constant updates to handle website changes. It's repetitive and tedious, but couldn't be automated due to the high data diversity and many edge cases. This required a combination of rule-based tools, developers, and constant maintenance.
We're now using LLMs to generate web scrapers and data transformation steps on the fly that adapt to website changes, automating the full process end-to-end.
I’d you’re interested in using one of the LLM-applications I have in prod, check out https://hex.tech/product/magic-ai/ It has a free limit every month to give it a try and see how you like it. If you have feedback after using it, we’re always very interested to hear from users.