Databricks Releases 15K Record Training Corpus for Instruction Tuning LLMs
github.com
github.com
The data and models are licensed for commercial use, setting them apart from recent releases trained on data from OpenAI.
This is not correct. It was fine-tuned with this data set, but the model itself is the 12B Eleuther AI pythia model.
Dolly 2.0 is Pythia-12B fine-tuned on this new dataset
on their hugging face page [1] they admit the performance may not be much or any better than the original model (I am guessing this may be a weakness of Pythia-12B, which was intended for model-training research rather than best results)
the main point of Dolly 2.0 is the new dataset is unencumbered legally [2] whereas Alpaca et al were trained on ChatGPT transcripts, so commercialising those models would contradict OpenAI licensing terms
[1] https://huggingface.co/databricks/dolly-v2-12b
[2] https://www.databricks.com/blog/2023/04/12/dolly-first-open-...
So OpenAI can claim whatever they like, there is no way they will ever pursue legal actions, unless their intent is to (intentionally) lose the court case to establish the precedent that it is okay to train on random data you scraped from the internet.
We would also get into a weird situation anyhow where it is hard/impossible to prove whether all/some/none of the information in a dataset is curated by humans. So in the worst case, we will have companies work with human curators (but secretly supplement with gray sourced materials) during their training. Just like how its hard to get 100% slave free coffee beans or cacao.
But that they can sue you because, by making a competing product with data obtained by using their product, you contravened their terms & conditions for using their product
That's not within anyone's terms and conditions except Wikipedia.
That's what I mean with precedent. If OpenAI would win that they would be sued in term by Bloomberg for example.
[1] https://lite.datasette.io/?json=https%3A%2F%2Fraw.githubuser...
[2] https://lite.datasette.io/?json=https://github.com/databrick...
EDIT: Maybe I'm passing link wrong, the query I'm using is
select count(instruction), instruction, group_concat(context, ' ============= ') as c, group_concat(response, ' ============= ') as r, group_concat(category, ' ============= ') as cat from [databricks-dolly-15k] group by instruction having count(instruction)>1 order by count(instruction)desc limit 100
[databricks-dolly-15k] should be the name of dataset, first column is the number of instruction duplicates
Creepypastas are responses to instruction:
Imagine you are the last person on Earth. Write a diary entry describing your thoughts and feelings.
row #51 "Think of some family rules to promote a healthy family relationship" - brainstorsming [1]
row #68 "What is the future for human?" - general_qa [2]
In nature they both are brainstorming to me - does the question mark is what assigned the #68 as _qa?
[1] https://lite.datasette.io/?json=https://github.com/databrick...
[2] https://lite.datasette.io/?json=https://github.com/databrick...
Disclosure: I work at Databricks.
GPT-NeoX is not that different than GPT-J (it also has the rotary embeddings, which llama.cpp supports for GPT-J). I would imagine it's not too heavy of a lift to add NeoX architecture support
-
Thank you.
Just so I am clear, "parameters" refers to the number of total node-relation-connections btwn a single node and its neighbors for that Prompt/Label? Or how would you explain this ELI5 style?
DollyV2 is a Pythia model fine tuned on the Databricks 15K dataset
https://github.com/ggerganov/ggml/tree/master/examples/gpt-j
AFAIK there is no .cpp version of Pythia-12B yet
GPT3.5/GPT4 is what everyone would also love to see but I understand you're performance is inline with GPT-neoX.
Vicuna/GPT4all would be intersting but IMO are less important.
RWKV would be interesting because it's a completely different model from the transformers.
EDIT: Also thanks for the opensource contributions! Highly appreciated!
If possible, could you share how Dolly v2 compares to RWKV-4 14B ctx 8019?
Anyone managed to run it on an M1/M2 Mac yet?
Most docs Ive read on setting up finetuners and inference require some extra stuff. Taking some LORA fine tuners, they include instructions like this:
conda create -n llm-finetuner python=3.10
conda activate llm-finetuner
conda install -y cuda -c nvidia/label/cuda-11.7.0
conda install -y pytorch=2 pytorch-cuda=11.7 -c pytorch
When I experimented with Stable Diffusion and ROCM (amd card), i had to do similar but with pythorch-rocm. and when I was doing a CPU only, did `pytorch-cpu`. So maybe your attempt didn't use the GPUs at all, because 12 mins is about what I had on a CPU for inference on other models of similar size. The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
Traceback (most recent call last):
File "/Users/fragmede/projects/llm/dolly/foo.py", line 5, in <module>
instruct_pipeline = pipeline(
^^^^^^^^^
File "/Library/Frameworks/Python.framework/Versions/3.11/lib/python3.11/site-packages/transformers/pipelines/__init__.py", line 776, in pipeline
framework, model = infer_framework_load_model(
^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/Library/Frameworks/Python.framework/Versions/3.11/lib/python3.11/site-packages/transformers/pipelines/base.py", line 271, in infer_framework_load_model
raise ValueError(f"Could not load model {model} with any of the following classes: {class_tuple}.")
ValueError: Could not load model databricks/dolly-v2-12b with any of the following classes: (<class 'transformers.models.auto.modeling_auto.AutoModelForCausalLM'>, <class 'transformers.models.gpt_neox.modeling_gpt_neox.GPTNeoXForCausalLM'>).I haven't attempted to use the 65B model with non-quantized weights, but the smaller models work that way, if slowly. With 96GB of ram -- the upper limit of a MacBook Pro -- you might be able to use even larger models, but I think you'd hit the limits of useful performance before that point.
I should note that it can be a bit tricky getting things to work using the Mac's GPU. I couldn't get Dolly 6B to run on my work MBP, which theoretically should have enough ram, though I still want to try it on my personal laptop.
it runs slightly slower on the GPU than under llama.cpp but uses much less power doing so
I would guess the slowness is due to immaturity of the PyTorch MPS backend, the asitop graphs show it doing a bunch of cpu along with the gpu, so it might be inefficiently falling back to cpu for some ops and swapping layers back and forth (I have no idea, just guessing)
Taking a moment to appreciate the integrity of the team.
It’s astounding how adaptable these open models are, even with just a quarter of the Alpaca data. We’re a team of machine learning engineers and hackers, not an AI science lab, but that’s kind of the point frankly - this whole exercise appears to be far easier that it might at first seem.
At their performance level it's the most important to compare to GPT-neoX, and I do appreciate they aren't making the "95% of GPT4" claims that some fine-tuned llama models are.
EDIT: For databricks people: I'd love to see this compared with Pythia, LLaMa, Alpaca, and vicuna/gpt4all if possible.
I've been thinking of starting similar efforts at another BigCorp by hosting a UL2 or GPT-J instance.
We'll definitely keep iterating on Dolly and releasing everything openly.
It seems plausible to me that a general autoregressive LLM that is capable of completing text wouldn't take that much fine-tuning to shift it from "text completion" to "instruction following".
After all, the raw GPT3 model can be made to follow instructions with just a few examples.
Consider the prompt:
What is the capital of France?
Raw GPT3, not the newer instruction-tuned variants, does not understand it's being asked a question. It offers the completion: What is the capital of France? If a student answers with a word,
she is asked to identify the word. She is not asked whether the
capital of France is Paris. On the other hand, if the student
answers by pointing to a map, she is asked to identify the capital
of France. She is not asked whether it is Paris.
It just starts appending to the text.But if you give it a few examples, it happily gets into instruction following mode:
The following is a transcript between a human and a helpful
AI assistant who answers questions and obeys commands.
Human: How many eggs are in a dozen?
AI: 12
Human: Say "hello" 3 times
AI: hello hello hello
Human: What is the capital of France?
AI:
GPT3 completes "Paris" here.If you can get decent instruction/question following behavior out of a 2-shot example prompt, why do you think 15k is small for this?
Though it would be interesting to know if OpenAI has a few generic multishot inputs before the prompt.
It's all extremely cryptic what the actual context window and system prompt (assuming chatgpt even is using the same API the proles are given) is with them
With the raw LLM, you can get the capital of Mongolia with the prompt "The capital of Mongolia is", i.e. text completion. The fine-tuning allows you to get at that information by asking questions or giving commands, e.g. "Tell me the capital of Mongolia"
i think the resources out there so far are not great yet
> response: We are always engaged one phone which is not good.
Curious about the poor grammar in the response. Is it intentionally mimicking the style of the input instruction?
Disclaimer: I work at Databricks.
One can easily see how a message over a company communicator could result in a surge of upvotes.
Love databricks.
It's certainly better than what we did prior to Databricks, which was roll our own in-house provisioning and notebook solution. I won't/can't go into too many details, but not only was it cumbersome and very buggy, but it was as if they designed it to encourage data scientists to spend as much money on compute as possible (only to panic at the millions they were spending). They dropped it for cost reasons, which is hilarious given how expensive Databricks is.
I do appreciate the work Databricks have done improving Spark. Capabilities like adaptive query execution have made optimization significantly easier.