You probably don’t need to fine-tune an LLM
tidepool.so
tidepool.so
Past few shot and RAG, you can overcome context window limits if you find ways to break a single request into many, each with specific context and then roll them up somehow. This can help get past context window limits.
Claude 2 has a large context window, but if you are actually giving that much in prompt examples, to cover tricky edge cases, I've found its better to break things down into multiple steps.
And if you can break things up that way, and costs isn't at issue, GPT-4, with lots of few shot examples, and chain of thought seems to give me the best results.
Or this is what I found writing a code translator for a language the LLM didn't know. I wrote it down in more details here:
For example, if I put instructions first, then a lot of context, it would forget the task, so instead put a small explanation, then context, then question, but still it would seem to sometimes loose the plot.
For feedback on prose though, I found it Claude very strong. It's nice to be able to give it half of a book, in text format and have it guide you to important parts and summarize sections. Some things, like what are the themes of this work, what's surprising finding, don't work well with vector DBs.
[1] https://www.pinecone.io/blog/why-use-retrieval-instead-of-la...
The tasks I'm doing might be uncommon. Here is a build script in some language you (the llm) understand, and I want to get it translated to a language you don't. And so the context is example conversions and documentation.
I've also had some luck with things that boil down to "Here is a very large style guide. now how would you improve this code?". Or "here is a number of examples of feedback on writing to conform to a style. now generate the same on this new input."
I found the large context windows and Claude to work quite well in those examples. But, if its possible, breaking it down into multiple steps with less context somehow and using GPT4 even better (though more work).
I imagine OpenAI is not far behind in expanding the context window of its models. The LLM companies have access to the same techniques and - in my estimation - are just choosing to focus on one aspect or another to address different market needs. For instance, Claude 2 clearly focuses on maximum context window size at the cost of speedy inference and, presumably, inference cost. By contrast, OpenAI seems to be focused on speed and low cost (GPT-3.5) and accuracy (GPT-4) rather than maximum token length.
https://www.anyscale.com/blog/fine-tuning-llama-2-a-comprehe...
We've got some additional resources for folks looking to better understand Retrieval Augmented Generation (RAG) and even see it in action - in this example we demonstrate a potentially very dangerous hallucination (that has to do with driving) and how to fix it using RAG: https://www.pinecone.io/learn/retrieval-augmented-generation...
If you're curious to actually try out the difference between an LLM without domain-specific context and an LLM that is using RAG, you can try our live demo here: https://pinecone-vercel-starter.vercel.app/
And if you'd like to fork and make your own tweaks to the above demo ^ chatbot, in order to, for example, swap in your own company logo and extend it for your purposes, you can find our Vercel template here: https://github.com/pinecone-io/pinecone-vercel-starter
In our opinion, RAG is indeed an effective technique partly because you don't need to be a machine learning expert in order to implement it in your Generative AI applications.
The future is likely > 90% of developers relying on the best frontier models and using the context to specialise, and 10% of specialised developers who have the expertise, budget, and time, customising LLMs for very specific use cases where there is no other option.
But I am very impressed with GPT-4's abilities in finding the correct answer just by reading those documents, so I think the only problem still enduring is the context window size.
That was my impression as well, which is why your comment was so interesting to me. Have you found tools/projects for the DAG approach that you'd recommend?
This is the easiest I found, on here too.
Does it not use any LLMs, or does it have its own, or how how does it work?
Thank you
Do you have any scripts you could share for the training/eval process? Would love to credit you in the post
With log file analysis as an example, training a model may increase the model's ability to deal with outliers, through writing regex which is placed in the indexing pipeline. In this use, tuning a prompt isn't going to help much, given the foundation model might have no idea how to parse a given field in a log line no matter how you put it to it in the prompt.
Tuning models also serves other purposes, such as removing guardrails introduced in the training data by others, and customizing the self referenced material the model "knows" about, such as its name, creators and the "personality" presented to the end user.
Finetuned Palm for Medicine and Finetuned Minerva for Math all perform a good deal worse than GPT-4.
A fine-tuned smaller model is by no means guaranteed to beat a larger more general one (though of course you may get acceptable performance).
And then the necessity of fine-tuning itself is called into question plenty with LLMs.
https://huggingface.co/papers/2308.00304
Fine-tuning even Lora on the open source models is nearly always better than these other approaches
A model's weights are inherently lossy and opaque. If a model asserts some fact, it is impossible to tell whether that fact was true or hallucinated just from the model, because the model has no notion of "fact" or "truth", it's just probabilities.
Generating an answer solely from model weights is like asking a random person to answer a question from memory. Sure, you might get the right answer, but there's no guarantee.
Using RAG is like handing them a book and asking them what the book says the answer is. With the benefit that LLMs can "read" much faster than a human.
My take is that LLMs are actually much better at "reading" than they are at "writing", and RAG plays to that strength.
They are certainly much faster at reading than at writing. In fact, they re-read the entire context window for every token they write!
Can you expand on that? I've not seen evidence of that myself yet, but maybe I haven't looked in the right places.
"Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning" https://arxiv.org/abs/2205.05638
Fine-tuning is not a solution for getting fresh data as you would with RAG, unless you are planning to run your entire fine-tuning suite for every new document. It can help improve accuracy a bit when you need to specialise the model for a very specific domain or modality. In practice, this is rare and unnecessary for most use cases building on LLMs.
The counter argument to your assertion (my own opinion, not Microsoft's) is that the reason you hear so much about fine-tuning from everyone other than Open AI / MS, is that they offer less capable models that can't reliably produce the same quality of results without fine-tuning.
This means you're easily looking at a 1 Million dollar project in order to be successful. Even once you're done, the odds of success are mixed - and Claude-3 may beat your fine-tuned model. These economics aren't hard for research shops, but startups are going to struggle with this approach.
const knownColumns = ['name', 'email', 'id']
template`
SELECT "${a('column name', {
sampler: bias.accept(oneOf(knownColumns))
})}"
FROM "table"
`
You're able to enforce, at the sampler level, that the output is one of the expected choices.2. Clearly specify the column names as part of the prompt and that there are no other columns.
3. Occasionally you may get an error that you have to feed back to GPT-4.
Future LLMs are likely to have much smaller size and have outside long term memory/training knowledge as well as work memory (a'la RAG approach).
- A RAG-empowered LLM can tell you where the knowledge used to answer a question came from.
What also matters is the size of the context window and how effectively the models can follow large amounts of instructions. So new models might change the advice again.
In humans, a person might say something like, "I'm not an expert, but X" and I think being able to default back to the underlying LLM would be useful.
But using the OpenAI API directly is not that complicated. And neither is using something like pg_vector or just a plain cosine similarity calculation from Stack Overflow.
I have been self hosting a 16K context size model and there is a lot you can do with 16K or larger context.
There are also great use cases for fine tuning. For example, if you are writing a chatbot for your company’s products, it might make sense to fine tune on product data and then RAG it with specific customer data when setting up a chat session.
“We’ll just fine-tune based on our (we think valuable + special) data”
Fine-tuning doesn’t enhance the model w/ new “knowledge” but a new narrowly-defined task
One other “cost” to consider is fine-tuning a 3rd-party model means if that foundation model changes or goes away that effort/cost needs to be repeated
OpenAI removed the base GPT-3.5 model a while ago and never made it available for GPT-4.
I’ve never seen it wander between genres, but it can tell pretty generic stories particularly if you’re pretty generic in your promoting.
If it is just text corpora, probably not worth it
I haven't heard of that, so genuinely curious. What I'm familiar with is guardrails applied on top of the LLM. For instance through prompting, managing the available data inside the vector database that could be used for context (if using RAG), or through something like NeMo[1].