Mistral Fine-Tune
github.com
github.com
For example, Bloomberg trained a GPT-3.5 class LLM on their financial data last year and soon after GPT-4-8k outperformed it on nearly all finance tasks.
We ended up focusing on having high-quality eval data and an architecture that makes switching to new models easy.
Also I think there’s a bunch of important techniques that just are t shared.
Here is an example of data prepared for fine tuning Llama for function calling...
https://huggingface.co/datasets/mzbac/function-calling-llama...
I'm unaware of any comprehensive guides - we're still in the wild west.
It's a narrow domain, but so is most of programming. I think if you're just training a general purpose LLM to be more inclined towards your data -- no, fine tuning is probably not very relevant. But if you're trying to solve a very specific yet fuzzy problem, and LLMs can get you _part_ of the way there, fine tuning is likely your best bet.
Many of the use cases I've seen for LLMs would actually be better with NLP.
Another use case for fine tuning we have as well is to reduce a 5-shot prompt that we have to run hundreds of times per request down to a “0-shot” (heavy emphasis on the double quotes.
Run five shot on gpt-4o a couple thousand times, then fine tune on cohere’s command-r or haiku or llama3 8b or whichever small but mighty llm. You can reduce costs by 99%, or somewhere in that ballpark, without really sacrificing quality on 99% of the queries.
There's a reference to the paper that describes the method at the bottom of the README: https://arxiv.org/pdf/2307.09702
Internal corporate data GPT4 was never exposed to?
(not mine)
In a strict sense, finetuning can add new knowledge but for that you need millions of tokens and multiple runs without using LoRA or Peft. For practical purposes, it does not.
But chatbots are only one single use-case. And broadly I think this pattern of LLM-as-store-of-knowledge is a bad one (of course, until ASI, and then it isn't).
That said, you absolutely can impart new knowledge through fine-tuning. Millions of tokens is a rather small hurdle to overcome. And if you're not retraining with original/general data, then your model will become very specialized and possibly overfit...which is not an issue in many instances, and may even be desirable.
I have non-English human data annotated in a format that was designed for a very specific health-related study. LLMs haver never seen these annotations, non-English LLMs are not a top priority for companies, and we can only use offline-first ones for data privacy reasons.
In this scenario fine-tuning a general purpose LM works wonders.
Other than that, fine-tunes don't really matter for us because not many people are rushing to beat the top models on (say) Georgian POS tagging or Urdu sentiment analysis.
As long as the model can turn language into a reasonable vector, we're happy with it.
We're considering finetuning for 2 reasons:
(1) the current system prompt with instructions is getting quite large* and we see more challenges for GPT to stick to the instructions. We could finetune with our historic summaries and a simplified prompt to slightly improve the performance of the summaries / lower token count (performance). The idea would be to then continue improving our system prompt again from that new starting point. Exploration still to be done though.
(2) we might be able to finetune a much smaller model to do what the larger model is currently doing (cost, sustainability)
* The larger instruction prompt is because there are lots of specific needs for the summaries (e.g. how to write the summary, what to exclude (health, etc.), what to include (actions taken), and we e.g. give examples of good vs bad summaries to improve performance.
I want it to be fixing knowledge based on my data and fine-tune it to be adjusted to my use case already.
I also want a.feedbackloop and can't start to send more and more context with the payload just because my feedback loop adds / finetunes data
RAG is a far better option for making a model work with your data/information.
Finetuning is better for making it output in a language, style, data format, programming language, etc.
in the current landscape, specialized, smaller systems will still be the most efficient way ahead in the near future.
If I wanna do this, what GPU would I need?? I have a 3060 TI (laptop one) and i9 with 16 gigs of ram. I don't have AWS quota, or quota in GCP either. have heard of paperspace, but I want to quickly get started with fine tuning Mistral, cause we are planning to use some of their models for an engagement I am working on.
They list hardware requirements by model and you can select your VRAM and system memory to filter available models.
Plus in a desktop, you have the ability to upgrade to faster GPU or even go for multiple GPUs.
Caveat: such machines, especially multi-GPU, are loud and produce enough heat to warm up a whole room quickly.
And cloud will likely be cheaper if you believe you will not spend more than 10% of the next couple of years running the GPU at full blast.
Power limiting Nvidia GPUs up to around 60% of max TDP can offer a substantial reduction to heat and noise with minimal performance loss, but agreed on everything else.
They also now do hourly rental.
The landscape is so fractured, I feel like I haven't even heard of most tools. I encountered Microsoft's olive the other day and it was completely new to me.
Even if in the repo it is stated that it is optimized for big models (needing A100/H100), I still feel this can help with smaller models more than large ones.
We can extend “If you build it, they will come” with “If you provide the tools, the will build”