Finetuning Large Language Models
magazine.sebastianraschka.com
magazine.sebastianraschka.com
If you want to be able to do Q&A against an existing corpus of documentation, can fine-tuning an LLM on that documentation get good results, or is that a waste of time compared to the trick where you search for relevant content and paste that into a prompt along with your question?
I see many people get excited about fine-tuning because they want to solve this problem.
The best answer I've seen so far is in https://github.com/openai/openai-cookbook/blob/main/examples...
> Although fine-tuning can feel like the more natural option—training on data is how GPT learned all of its other knowledge, after all—we generally do not recommend it as a way to teach the model knowledge. Fine-tuning is better suited to teaching specialized tasks or styles, and is less reliable for factual recall. [...] In contrast, message inputs are like short-term memory. When you insert knowledge into a message, it’s like taking an exam with open notes. With notes in hand, the model is more likely to arrive at correct answers.
Ideally, we include documents that are relevant to the topics at hand and eliminate documents that aren't. Using prompt chaining, or an iterative process of labeling and extracting features, we may be able to increase the efficiency of the final prompt.
Obviously this increases costs, but so does repeated fine-tuning.
It seems pretty expensive because you may be pasting a lot of context into each query. If you allow the user to have follow up queries and you want to retain context of the conversation that seems expensive too. But it does seem like it should give the best results for Q&A as long as question is directly answered in your data somewhere.
Github Copilot seems to be the most effective code assistant currently. It seems to use many heuristics for figuring out relevant snippets to include in prompts, like computing Jaccard similarity of windows of the last 20 opened files. It also tries some tree-sitter cleverness for some languages, but when I snoop on the HTTP traffic it seems to almost always just give up and only include the rest of a file as context.
I have wondered whether a model fine-tuned on my own code would do much better, and be simpler. But perhaps building embeddings and searching them (like in the article you linked) would be superior.
Code assistants need to be super low latency though which maybe complicates things too.
I honestly think that if you could have all your private code indexed and accessible, this would be a game changer as it has way better context.
My open questions though are:
- does the tokenizer for the index need to be the same as the one used by the LLM?
- is running a local LLM feasible, or should I just assume prompts must sent over the network? This would affect latency and privacy a lot.
- is fine tuning helpful? Or necessary, even, to ensure the LLM only provides source code, and only in the target language?
Also possible when using the Vim copilot plugin (or its emacs port) which is pretty informative - you can see a lot of the extra tricks the VSCode implementation uses.
Does anybody know if this is already available on Hugging Face or somewhere else?
I'm using TypeScript, so my idea would be to finetune on all used packages, maybe even all of the documentation as well.
I've created a small app where I create a prompt for an eslint error and let ChatGPT come up with a git diff. I manually feed that back to my application, it eslints and compiles and if there are any errors, a new prompt is generated. This works quite well.
GPT-4 is not available for me in the OpenAI API, so it's quite cumbersome.
*The merged model contains the LoRA. Applying the LoRA over a "raw" llama model uses more VRAM and does not allow for full context in 24GB of VRAM.
I'm also super keen for the 32k API limit for OpenAI, that's going to be great.
But, as far as I know, OpenAI will not help you fine-tune, they will not run a fine-tuned model for you, and they would probably prefer that you use their models over open models that can be fine tuned.
(None of this is to say that fine tuning is better. I’m just saying that OpenAI has a strong commercial bias.)
I like how the fine-tuning page talks about how fine tuning is supposedly better, though.
Although a lot of people are working on hallucination reduction, I think we're a long way off from that in the general case. So having the ability to point to a real piece of data, outside the model, is important for applications where accuracy matters.
https://community.openai.com/t/fine-tuning-myths-openai-docu...
Bottom line is that fine-tuning does not seem to be a feasible option for adding new knowledge to a model for question answering.
https://adapterhub.ml/blog/2022/03/adapter-transformers-v3-u...
"OpenAI Q&A: Finetuning GPT-3 vs. Semantic Search - which to use, when, and why?"
Supabase Clippy was the first docs site to ship this experience to production as far as I can tell: https://supabase.com/blog/chatgpt-supabase-docs
I believe they called it "context injection" and I have been following suit in my own writing on the topic.
I am prototyping experiences like Supabase Clippy and am also very interested in fine-tuning for docs Q&A. But my main question is: what exactly would the fine-tuning inputs and outputs look like for docs Q&A?
Edit: For Q&A the question is the input and the answer is the desired output? Is that right?
A more general comment about fine-tuning for docs from my blog:
> AI is all about prediction. Given this temperature, this wind, this day of the year, what is the chance of rain? Temperature, wind, and date are your inputs. Chance of rain is your desired output. Now, try to apply this same type of thinking towards documentation. What are your inputs? What’s your output? The page title and code block could be your inputs. Whether or not the code builds could be your output. Or maybe the code block should be the output? This is why I keep saying that applying fine-tuning to docs is tricky. What are the inputs and outputs?
https://technicalwriting.tools/posts/ten-principles-response...
(I am an AI n00b and have not looked deeply into how fine-tuning works but it's high on my list to experiment with OpenAI's fine-tuning API. Please LMK if I am getting any fundamentals wrong.)
https://help.openai.com/en/articles/7127982-can-i-fine-tune-...
If you want exact detailed recall, then using a framework that provides search and recall (by embeddings or otherwise) is probably always going to beat fine tuning, but also remember, it doesn’t have to be either-or.
I mean, if you want a person to handle Q&A on a corpus, is it better for them to have studied it, or have direct access to the corpus with an appropriate index? The answer is clearly that its better if they’ve trained and have access to the corpus, and while LLMs aren’t the same as people, I think the answer for them is the same here.
Our use case was/is a text-to-sql bot that would provide domain-specific output. Our domain is very complex so any sort of off-the-shelf AI SQL helper is a joke to us at best. My thinking goes something like "If your schema is so simple that you don't need to think about FTing the model, why do you need its help writing SQL in the first place? Your answers are probably on google somewhere."
The perspective we have now is that the LLM is a probabilistic text transformation engine that must always be contextualized by some external means. Clearly, if we could just include all of our context & it fits & it's cheap, we should just do this. But, the reason anyone is thinking about FT in the first place is because you can't fit the whole damn business in the prompt (or because it's too expensive to do it at scale).
The approach we are looking into now is a classification front end that attempts to discover the relevant business context being referred to by the user's initial prompt. Once we detect our classes, we look up the boilerplate text and incorporate it into the final prompt.
So, instead of trying to align the Jupiter-scale model to your business needs, leave it alone and build a smaller adapter layer that can work with any unmodified LLM.
All I can really add - Binary classification is like a super power once you understand the statistics around it. If you can find clever ways to combine multiple binary classifiers, you can quickly narrow down relevant context. You can also use the statistics to do things like determine if a query is too vague to be serviced by an LLM backend. You can also answer questions like "Why didn't we consider a table or join when writing the user's query?".
I'd definitely like to see how others are doing it.
[1] https://huggingface.co/datasets/b-mc2/sql-create-context
Take GPT-4, for example. GPT-4's zero-shot prompt performance is excellent. It capably handles tons of tasks—even in specialized domains such as medicine and law. (Note: I'm a doctor and lawyer both.) I can see future LLMs being even more capable.
In fact, I see LLMs as platforms rather than products. Take smartphones: you have just two dominant platforms, iOS and Android. An entire ecosystem of apps runs on them. A software-only example might be web browsers: Chrome (and chromium browsers) and Firefox. These are the platforms on which millions of web applications run. I see LLMs ending up this way: a few platform providers, and a much vaster ecosystem of "apps" built on them.
This also implies something else: the "fine-tuning data" moat you thought your organization might have had might not be a moat at all. "Foundation" LLMs will become so good that everyone else will be able to do what you're doing. You'll have to compete the old-fashioned way: by building a better product, and crushing sales and marketing.
Edit: This affects every use case where each prompt-completion pair provides new and essential information that the model needs to take into account.
Now, the cost of GPT-4 is an issue. I hope that open-source models such those released by Stability AI reach the same level of quality as (at least) GPT-3.5-turbo soon. That would be a game changer.
I tried out da-vinci, which also works fine, but models below da-vinci don't work well for my use case.
For extracting what it was if it's in the document, you can use the indexing approach described in the article. For predicting what it should be, you'd probably be better off using llm as an extraction mechanism and then using structured data models to classify discharge disposition
A lot of factors go into predicting the post-acute care route that results in the best outcome. You are correct that it would be possible to create features using an LLM, but that is a very difficult problem compared to simply treating the problem as a sequence classification.
But yeah my money is on context windows getting bigger and frameworks smoothing out how to do the above automatically, So it feels like a point in time optimization right now That will be tech debt by the end of the year
> That being said, you can run LLMs on-prem.
how is this relevant at all? it's just dodging the question, since the claim was about gpt4 32k tokens and not the crop of generic models that we have now availabe on prem (and which don't support 32k token anyway)
"enhancements to automatically draft message responses"
"Another solution will bring natural language queries to SlicerDicer"
neither of which needs 32k token nor to see the patient records
also, foot note 1
"users of Azure and Azure OpenAI Service are responsible for ensuring the regulatory compliance of their use"
there is zero claim of actual compliances of these services for handling sensitive or regulated data.
now I understand the enthusiasm, but at least don't waste people time with sources that are, at best, tangential, and don't provide any substance to the discussion
"The second use will bring natural language queries and "data analysis" to SlicerDicer, which is Epic's data-exploration tool that allows searches across large numbers of patients to identify trends that could be useful for making new discoveries or for financial reasons. According to Microsoft, that will help "clinical leaders explore data in a conversational and intuitive way.""
Do you not understand what natural language queries on external data by access of/enabled by GPT-4 means/entails ?
The only thing worse than being condescending is being condescending and wrong.
that means natural language will understand user request and write queries for SlicerDicer data backend, not that data is sent to the llm for analysis.
it writes queries, don't receive data.
Not just opinion, GPT-3.5 Turbo is demonstrably worse at just about everything except providing so called "safe" and "aligned" replies to questions.
Translate the following German sentences into English: Example 1: German: "Ich liebe Eis." English: "I love ice cream." Example 2: German: "Draußen ist es stürmisch und regnerisch" English: "It's stormy and rainy outside Translate this sentence: German: "Wo ist die naechste Supermarkt?"
Why is this needed when saying "Translate this to German: Hello how are you" just works out of the box?
In-context learning with few-shot exemplars are needed for things like encouraging responses that adhere to a certain JSON data type.
I imagine down the road we may learn what to tune into models dynamically and how deep it needs to go. Or maybe the context window becomes a sort of rapidly trained thin NN layer on top like we’ve seen in diffusion models
https://www.nuget.org/packages/OpenAILib
Instead of having to dealing with all of the various end points, creating files in a certain format (and then use the same conventions when using the model), you can simply create a fine tune and use it in completion requests.
Has anyone on HN tried replacing systems like Elasticsearch with an LLM based index? Curious.
One of the systems at my startup is an elasticsearch based search of a large corpus of structured data that contains larger text fields.
Cost is definitely an issue.
I've looked at haystack a little bit and was wondering how involved it would be to set up for a research spike.