Patterns for building LLM-based systems and products
eugeneyan.com
eugeneyan.com
I love articles like these, and how they are able to bring me up to speed (at least to some degree) on the "new paradigm" that is AI/LLM.
As a coder I cannot say what the future will look like (binary views) but I can easily believe that in the future we will have MORE AI/LLM and not LESS AI/LLM thus getting up to speed (at least on the acronyms and core theory and concepts) is well worthwhile.
Very Good Article !
And neither of the articles even mentions the other.
Sometimes the terms and metrics are borrowed from the (paper) authors domain, other times they're just coined on the fly - if there are no good analogies, or there are too many overlapping or closely related terms. It's like in math where you have a double-digit number of notations for the dot product - all depending on what sub-field of math you're working on.
I understand WHY it is that way, but it's super frustrating - because you end up spending time to look up what the notation / terms mean, often with no real clear answer.
That's when you know it's going to be amazing. Single best narrative-form overview of the current state of integrating LLM's into applications and the challenges encountered I've read so far. This is fantastic and must have required an incredible amount of work. Massive kudos to the author.
I recently trialed an AI Therapy Assistant service. If I stayed on topic, then it stayed on topic. If I asked it to generate poems or code samples, it happily did that too.
It felt like they rushed it out without even considering that someone might ask it non-therapy related questions.
I’ve definitely talked poetry and writing with a therapist, and while I’ve never had my therapist provide code, we’ve definitely talked tech in great detail.
Maybe those therapists were intentionally making me comfortable by engaging with shared interests. And the LLM isn’t being intentional about it, but I’m not convinced that a therapist is ineffective if they fail to stay ‘on topic’ when directed off topic by their patient.
"Is this question an appropriate question to ask a Therapy Assistant please respond with a single word Yes or No"
or something like that. will it be perfect probably not. but I mean it's only mental health what could go wrong...
Most corpos couldn't give a rat's ass about it, it's just the fancy new toy on the block that's saturating everyone's newsfeeds so they have to jump on it lest they be left in the dust by the competition who are doing the exact same shit, aka calling the "Open"AI APIs and pretending they're doing something groundbreaking.
We got interrupted mid-sprint, mid-epic to make some shitty wrapper around their APIs. I suspect the overwhelming majority of companies with fancy new "AI" features are doing the exact same shit
Evals, RAG, Guardrails often times require recursive calls to LLM's or other fine-tuned systems which are based on LLM's.
I would like to see LLM's and models condensed and bundled up into more singular task trained models - much more beneficial versus system design on using LLM's for applications.
This seems like we are applying traditional system design patterns for using LLM's in practice in apps.
Pretty telling about where we are in the evolution of these systems.
- any references for how this hybrid retrieval is done?
You can do approximate KNN search with ES by adding a setting on the index that enables KNN and then creating mappings for your embedding objects that defines their vector length. Then index your data as you normally would plus embeddings.
Once you have those in place you can construct your query and include the embedding similarity in how the query gets scored. When a query is submitted you embed it and pass into your ES query the embedding as well as the original query. ES will combine all of these elements together to score the results.
Basically, get the BM25 results and normalize them to be between 0 and 1, then take the (potentially weighted) average of them and the cosine similarity results (already between 0 and 1) to get the final ranking.
TL;DR - doing hybrid vector + keyword search provides more relevant results for text searches than vector search alone. And using sparse vector embeddings for the “keyword” part provides even more relevant results than using BM25.
I'd be more interested in the sales and marketing patterns being employed to hawk the same rebranded wrappers over and over. Ultimately, that's what's really going to contribute most to the success of all these startups.
Is a navigator a trivial wrapper? Is a sophisticated set of fine tuning and prompt Management a trivial wrapper?
What is a trivial wrapper?
I’ve heard Dropbox be called a trivial wrapper but at the end of the day it’s a successful company and the naysayers here are… not.
You'd either need access to the model weights or a fine-tuning API.
Then depending on which fine-tuning approach you want to use, the user data you need to collect will be different: RLHF requires multiple outputs to a single query vs instruction fine-tuning where you need great input-output pairs to train on. You could ask the user's feedback after running the LLM to pick out good training data.
where's the dual LLM pattern?
Practically speaking the starting point should be things like the APIs such as OpenAI or open source frameworks and software. For example, llama_index https://github.com/jerryjliu/llama_index. You can use something like that or another GitHub repo built with it to create a customized chatbot application in a few minutes or a few days. (It should not take two weeks and $15,000).
It would be good to see something detailed that demonstrates an actual use case for fine tuning. Also, I don't believe that the academic tests are appropriate in that case. If you really were dead set on avoiding a leading edge closed LLM, and doing actual fine-tuning, you would want a person to look at the outputs and judge them in their specific context such as handling customer support requests for that system.
100% this web page (or similar) will be used to basically scam clients into overpaying for simple wrappers around llama_index or LangChain etc. Some people will spend a week wasting their time trying to fine-tune an open source LLM on some wholly inadequate dataset before realizing they can use OpenAI and something from github. But most will not admit that.
Sure, a few people doing basically research projects for a large company or university will find some of the information useful. But realistically, probably not so much if they have to actually deliver a working system in a reasonable amount of time that would justify the business expense.
I'll grant you that the way the author presents his ideas seems a bit academic, but I assure you all of this information is just the immediate things you run into as a software engineer trying to integrate LLM's into your systems. No one is jumping to finetuning before they've even tried GPT4 on their problem.
Source: finished a chat to your documents project a month ago
Finetuning is not that needed in my experience.
rsync
On the other hand, directly using an off-the-shelf model, even the best ones, may not meet your performance requirements.
That’s where fine-tuning an open LLM is necessary.
Did you miss the NFT train? Have you ever asked yourself if this is what you should be doing with your life?
Just speaking as a guy who actually writes logic and code, rather than like, coming up with incantations and selling horseshit.
This will be on youtube as part of various "Top 5 tips to run AI in your app" in about 2 to 3 months.
And yes right now the challenge is figuring out how to productize these LLM things. The researchers are off figuring out what comes after LLMs, we're over here figuring out what to do with these things and how.
If you treat coding like learning the cheats of a videogame, you're not a coder, you're not a hacker, you're just a gamer and a fanboi. A consumer of whatever you're given.