Fine-Tuning Llama-2: A Comprehensive Case Study for Tailoring Custom Models
anyscale.com
anyscale.com
Fine-tuning Llama stream: https://www.youtube.com/watch?v=TYgtG2Th6fI&t=2282s
I have a couple more one where I do a QLoRa fine tuning session and explain the concepts as a personally self taught engineer (software engineer of 8 years moving into ML recently)
QloRa fine-tuning stream: https://www.youtube.com/watch?v=LitybCiLhSc&t=4584s
Overall I'm trying to breakdown how I'm approaching a lot of my personal projects and my current AI driven startup. Want to make this information as accessible as possible. Also have a series where I'm fine-tuning a model to be the smallest webdev llm as possible which seems like people are liking. Only been streaming for about a month and plenty more to come.
Ask me any question about the stream and fine-tuning llama!
If you have a 24GB VRAM GPU like a RTX 3090/4090, you can Qlora finetune a 13B or even a 30B model (in a few hours).
How does segmenting fine tuning models make sense? Do I need a terraform LLM, a SQL LLM, and a python LLM, or can I just use a “code” LLM?
In your example, you would fine tune the model to train it to code in a language it hasn't seen before, RAG will not really help with that.
For example, changing the way in which it responds could be:
- debate me
- brainstorm
- be sarcastic
Which also seems like something that could be accomplished with a system prompt or few shot examples, so I'm not sure when SFT is the more appropriate approach or what the tradeoffs are.Alternatively, gaining new knowledge would be training it on a dataset of e.g. sports trivia to make it highly effective at answering those types of questions.
P.S. nice username... Irving Fisher would approve.
Overall it depends on whether or not you can turn your data into a fine-tuning data and if you can find a low parameter (enough) model that can use your found contexts as input to host yourself of use inference endpoints. Hosting an LLM is actually not easy and I'm finding in the field working an information retrieval business OpenAI isn't terrible compared to costs of having a GPUs for your users across the world.
Everybody new to this field thinks that he needs finetuning to teach the LLM of new facts. I made the same mistake initially, later I published a slightly ranty post on that: https://zzbbyy.substack.com/p/why-you-need-rag-not-finetunin...
too much implementation detail required make it inaccessible for any non-significant use case. i imagine privateGpt will get there slowly
That's the problem I've been facing with Llama 2 as well. It's almost impossible to have it just output the desired text. It will always add something before and after its response. Does anyone know if there's any prompt technique to fix this problem?
airoboros supports the PLAINFORMAT token "to avoid backticks, explanations, etc. and just print the code".
https://huggingface.co/TheBloke/airoboros-l2-70B-GPT4-2.0-GG...
I wonder if LLMs will have less reasoning power if they simply return the output. AFAIK, they think by writing their thoughts. So forcing an LLM to just return the goddamn code might limit its reasoning skills, leading to poor code. Is that true?
In practice I haven't seen it make too much of a difference with GPT. The model can still use comments to express itself.
For non coding tasks, adding "Think step by step" makes a huge difference (versus YOLOing a single word reply).
Yes you're right. I'm mostly concerned with the text that actually "computes" something before the actual code begins. Niceties like "sure! happy to help" don't compute anything.
CoT indeed works. Now I've seem people take it to the extreme by having tree of thoughts, forest of thoughts, etc. but I'm not sure how much "reasoning" we can extract from a model that is obviously limited in terms of knowledge and intelligence. CoT already gets us to 80% of the way. With some tweaks it can get even better.
I've also seen simulation methods where GPT "agents" talk to each other to form better ideas about a subject. But then again, it's like trying to achieve perpetual motion in physics. One can't get more intelligence from a system than one puts in the system.
Not necessarily the same thing, as you're still putting in more processing power/checking more possible paths. Its kinda like simulated annealing, sure the system is dumb, but as long as checking if you have a correct answer is cheap, it still narrows down the search space a lot.
Yeah I get that. We assume there's X amount of intelligence in the LLM and try different paths to tap on that potential. The more paths are simulated, the closer we get to the LLM's intelligence asymptote. But then that's it—we can't go any further.
There's no reason to handle the LLM side of things, unless you want to try and optimize the amount of tokens which are code vs comments vs explanations and such. (Though you could also just start a new context window with only your code or such)
Really like the evaluation methodology, and seems well-written as well.
I don't think it should be something brushed on the side to be tried out later..
From a practical perspective, unless cost is really immaterial, I think most will end up starting with Lora, especially for 13b or 70b models.. you could do 10 fine-tuning runs for the cost of a few full fine-tunings.
But it's still all witchcraft to me to some degree, and I'd probably try full and Lora.
"For the 7B and 13B models, we used 16xA10Gs, and for the 70B model, we used 32xA10Gs (across 4x g5.48xlarge instances). When using Ray, there's no need to secure A100s to perform full-parameter fine-tuning on these models! The process is simply repeated for each task. Figures below show an example run based on a context length of 512, with a total of 3.7M effective tokens per epoch on GSM8k dataset.
We ran the training for a maximum of 10 epochs and selected the best checkpoint according to the minimum perplexity score on the validation set."
Has anyone taken them to court about this? Do we all just decide it's not fair and ignore it?
This blog seems to got good attention :) So we definitely plan to add it to Ray Summit https://raysummit.anyscale.com/agenda
Please comment on this thread if you have ideas of what kind of content you want to see more at Ray Summit
> At least 1xg5.16xlarge for head-node and 15xg5.4xlarge for worker nodes for both 7B and 13B
For the uninitiated, anyone have an idea how much this would cost on AWS?
g5.4xlarge - $1.6240/hour
You're looking at about $30/hour to run this in us-east-1.
https://instances.vantage.sh/?selected=g5.16xlarge,g5.4xlarg...