Optimizing LLMs from a Dataset Perspective
sebastianraschka.com
sebastianraschka.com
1) Train an initial big model on everything you can get, yielding a capable but tainted-in-some-jurisdictions model. Keep that model private.
2) Use the big tainted model to narrow or distill the source data. One way is by identifying the document subset that can be used freely (old public domain works, user generated content uploaded to your own service that users already assented to your own company's ToS on, government documents, things with unrestricted Creative Commons licensing...) The other way is by using it to build "just the facts" distillations from restrictively licensed material.
3) Train an untainted model using just the factual distillations and/or the permissively licensed material.
Phi-1 and therefore phi-1.5 are partially trained on gpt3.5 generated synthetic textbooks.
See https://research.ibm.com/blog/retrieval-augmented-generation...
I want to try fine-tuning to machine translate to and from a fairly niche language (https://en.wikipedia.org/wiki/S'gaw_Karen_language). How much text would I need, and what format would be ideal?
I have a number of book length texts, most only in the target language, and a few bilingual or multilingual. For the bilingual and multilingual texts, I can script out probably several thousand pairs of "translate the following text from <source_lang> to <target_lang>: <source_lang_text> <target_lang_text>". Do I need to vary the prompt and format, or can I expect the LLM to generalize to different translation requests? Is there value in repeating the material in different lengths? One set of sentence lengths, another paragraph, and another page or chapter length? Also what should be done with the monolingual texts, just ignore them?
It might be beneficial to start your dataset at the key (word) level, generate some embeddings of the key pair in the source and target and stash them, then do the same for sentence level and just for fun, paragraph level. (I believe you could get enough context from the sentence level as a paragraph is just a group of sentences but it would still be interesting to generate paragraph level key pairs I think).
From there you’d have a set of embeddings of each word src:tgt that also has context of how it fits in a sentence level and paragraph level with the respective nuances of each language.
Once you have that dataset then you can augment your data with prompts like you’re using but also including some contextual references of word pairs, and sentence pairs in your prompt which should corner the LLM into the right path.
Edit: not an expert so will heed if someone smarter comes along.
As noted below, extracting words or keyterms would maybe be a good idea, as they could be included in the training set.
The training set would the be comprised of the prompt, the translation, and keyterms. As you will want to vet the generated texts anyway, you could then decide if the foundational model was working enough. You could also try to run the largest "open" model you could find on the prompts, to see if those needed training as well. There are many different Llama models trained on HuggingFace for language pairs, so see if your languages are already built and test those.
I'm building a simple, Open Source ML pipeline manager at https://ai.featurebase.com/. I'd be down to help you with this!
I need to train more models to see if this is an accurate claim, but I've been finishing up the storage layer and haven't gotten to that yet.
Taking it a step further, I would include in the demonstration a test harness set up with a test suite to prove the proposed implementation.
I would go through each demonstration which a fixed set of criteria measuring only passing tests but ones that also show a level of complexity and usefulness.
Why? I was looking through CodeLlamas demonstration data for fine tuning and saw answers that were not even checked for correctness or usefulness.
Here's one library to do this https://github.com/guidance-ai/guidance
Keyterms can be used in the prompt to drive the LLM to better grounded responses as well as helping locate relevant embeddings for RAG, when vector search isn't enough.
Another consideration is writing code for processing things that look similar. For example, one might have the LLM write regex code which is then tested to work and put into production in a pipeline to parse log files, or write SQL off conversational queries, which are then run against a database.