State-of-the-art open-source chatbot, Vicuna-13B, just released model weights
twitter.com
twitter.com
> We release Vicuna weights as delta weights to comply with the LLaMA model
> license. You can add our delta to the original LLaMA weights to obtain
> the Vicuna weights.
Edit: took me a while to find it, here's a direct link to the delta weights: https://huggingface.co/lmsys/vicuna-13b-delta-v0Edit, later: I found some instructive pages on how to use the vicuna weights with llama.cpp (https://lmsysvicuna.miraheze.org/wiki/How_to_use_Vicuna#Use_...) and pre-made ggml format compatible 4-bit quantized vicuna weights, https://huggingface.co/eachadea/ggml-vicuna-13b-4bit/tree/ma... (8GB ready to go, no 60+GB RAM steps needed)
``` ValueError: Tokenizer class LLaMATokenizer does not exist or is not currently imported. ```
This and using conda in wsl2, instead on bare windows
(I know a vicuna is a llama like animal.)
We’re fighting back against the DMCA requests on the basis that NN weights aren’t copyrightable. This thread has details: https://news.ycombinator.com/item?id=35393782
I don't think you have to worry about Facebook going after you. The worst that will happen is that they issue a DMCA, in which case your project gets knocked offline. I don’t think they’ll be going the RIAA route of suing individual hackers.
The DMCAs were also launched by a third party law firm, not Meta themselves, so there’s a bit of “left hand doesn’t know what the right hand is doing” in all of this.
I’ll keep everyone updated. For now, hack freely.
creating sentient life?
on what legal theory or precedence makes this true?
IMHO, the weights are akin to the list of telephone numbers in a directory - which is definitely not copyrightable; only the layouts and expressive portion of a phone directory is copyrightable.
So to make the weights copyrightable, it needs to be argued that the 'layout' of the weight is a creative expression, rather than a 'fact'. But the weights are matrices , which is not expressive or creative. Someone else could derive this exact same set of weights from scratch via the same algorithmic procedure, and therefore, these weights cannot be a creative expression.
Firstly, weights are not merely a collection of facts like a telephone book is. If two companies train two LLMs they'll get different weights every time. The weights are fundamentally derived from the creative choices they make around hyperparameter selection, training data choices, algorithmic tweaks etc.
Secondly, weights can be considered software and software is copyrightable. You might consider it obvious that weights are not software, but to argue this you'd need an argument that also generalizes to other things that are commonly considered to be copyrightable like compiled binaries, application data files and so on. You'd also need to tackle the argument that weights have no value without the software that uses them (and thus are an extension of that software).
Finally, there's the practical argument. Weights should be copyrightable because they cost a lot of money to produce, society benefits from having large models exist, and this requires them to be treated as the private property of whoever creates them. This latter one should in theory more be a political matter, but copyright law is vague enough that it can come down to a social decision by judges.
Recipes, famously, are almost but not quite copyrightable | patentable.
eg:
https://copyrightalliance.org/are-recipes-cookbooks-protecte...
https://etheringtons.com.au/are-recipes-protected-by-copyrig...
I would contest the analogy, but even if we accept it, it's still not clear whether phone directories (or other compilation of factual data) are definitely not copyrightable. The position is clear in the US, but in the UK and presumably other jurisdictions, I wouldn't be so sure.
You could claim we're just talking about US law here, but if you release something on github/huggingface without geo-restrictions, and your company does business in Europe, you might not only have to comply with US law...
eg. https://www.jstor.org/stable/24866738 , eg. https://books.google.com.hk/books?id=wHJBemWuPT4C&pg=PA114&l...
What do you mean? IP law is overwhelmingly an impediment to progress; innovation happens faster when people are free to build on existing weights.
Removing IP protection would make it a lot easier to innovate at this level, but it would reduce the amount of money flowing into getting us to the next level.
Somehow, though, I doubt they'll let the golden goose slip through their fingers, no matter what happens.
This won't happen to GPT any time soon so they are safe, copyright or not.
There's more detail about the upsides/downsides in this thread: https://twitter.com/theshawwn/status/1641804013791215619
1- to have large corporations and people with privileged access to them have these models exclusively and have them collaborate as a clique
2- to have those models openly released to everybody, or de-facto released to everybody as they leak in short order
3- to have the people who think releasing models is a bad thing simply not release them and work alone in their proprietary solutions, as the smaller companies and hobbyists do collaborate
i say let them have a go at number 3 and see how that works for them - shades of "Microsoft Network" vs Internet all over again
one has to wonder how did they implement the storage of those deltas to require that sort of RAM
Though I expect somebody to write a patch to make this more accessible to people on laptops.
Edit: Also, judging by a comment from the team in the GitHub repository (https://github.com/lm-sys/FastChat/issues/86#issuecomment-14...), they seem to at least hint about been in contact with the llama team.
Vicuna-13B: The legality of a man marrying his widow's sister (widow marriage) varies depending on the country and cultural context. In some societies, widow marriage was traditionally practiced or tolerated.
GPT-4: It is not possible for a man to marry his widow's sister, as the term "widow" refers to a deceased man's wife. Therefore, if a man has a widow, he is deceased and therefore cannot marry anyone.
Koala: The legality of a man marrying his widow's sister varies by jurisdiction. In some countries, such as the United States, it is generally legal for a man to marry his widow's sister. In other countries, such as the United Kingdom, it is illegal for a man to marry his widow's sister because it is considered incest. It is always important to check the laws and regulations in the jurisdiction in which the marriage will take place to ensure that it is legal.
You'd probably need to come up with a new one now though, or confirm knowledge cutoff for the next evaluation :p
It's pretty awesome to realize that from now onward my computers are going to be able to help catch more and more of the holes that clearly exist in my cognition.
as if it were a common term.
Doesn't make Vicuna less impressive, it comes pretty close to Chat-GPT in many regards. And I like that trick question.
I've recently opened a GitHub repository which includes information for both AI model series[0] and frontends you can use to run them[1]. I've wrote a Reddit post beforehand that's messier, but a lot more technical[2].
I try to keep them as up-to-date as possible, but I might've missed something or my info may not be completely accurate. It's mostly to help get people's feet wet.
[0] - https://github.com/Crataco/ai-guide/blob/main/guide/models.m...
[1] - https://github.com/Crataco/ai-guide/blob/main/guide/frontend...
[2] - https://old.reddit.com/user/Crataco/comments/zuowi9/opensour...
these could be useful:
https://github.com/Crataco/ai-guide/blob/main/guide/models.m... -> https://old.reddit.com/user/Crataco/comments/zuowi9/opensour...
https://github.com/cocktailpeanut/dalai
the 4-bit quantized version of LLaMA 13B runs on my laptop without a dedicated GPU and I guess the same would apply to quantized vicuna 13B but I haven't tried that yet (converted as in this link but for 13B instead of 7B https://github.com/ggerganov/llama.cpp#usage )
GPT4All Lora's also works, perhaps the most compelling results I've got yet in my local computer - I have to try quantized Vicuna to see how that one goes, but processing the files to get a 4bit quantized version will take many hours so I'm a bit hesitant
PS: converting 13B Llama took my laptop's i7 around 20 hours and required a large swap file on top of its 16GB of RAM
feel free to answer back if you're trying any of these things this week (later I might lose track)
On that note, why is any RAM needed? Can't the files be loaded and diffed chunk by chunk?
Edit: The docs for running Koala (a similar model) locally say this (about converting LLaMA to Koala):
>To facilitate training very large language models that does not fit into the main memory of a single machine, EasyLM adopt a streaming format of model checkpoint. The streaming checkpointing format is implemented in checkpoint.py. During checkpointing, the StreamingCheckpointer simply flatten a nested state dictionary into a single level dictionary, and stream the key, value pairs to a file one by one using messagepack. Because it streams the tensors one by one, the checkpointer only needs to gather one tensor from the distributed accelerators to the main memory at a time, hence saving a lot of memory.
https://github.com/young-geng/EasyLM/blob/main/docs/checkpoi...
https://github.com/young-geng/EasyLM/blob/main/docs/koala.md
Presumably the same technique can be used with Vicuna.
I tried a few from https://www.jailbreakchat.com/ and it refused them all. Interesting.
Last fall it seemed that all the stars have aligned. The crypto winter and Ethereum switching to proof of stake meant that GPU prices fell to a reasonable level, I knew i would have a bit of a time to play some game during the holidays and as soon as Stable Diffusion was first posted on hacker news I knew that that's my excuse and my sign.
So far I think I have spent more time tinkering with the 20 python environments I have[0] for all the ML projects than playing RDR2.
Alpaca-30B is much better, it will even tell you how to build a nuclear weapon (incorrectly, of course, it’s not that smart).
I am waiting for Coati13B weights, these should work great.
I'll be busy next few days. Heck yeah.
Edit: just finished the conversion of Vicuna myself now and been doing some light testing, seems to work in ~80% of the cases for it, not as high success-rate as with GPT for sure. Probably there is a better way of structuring the prompt for Vicuna.
> write a function in JavaScript that turns a JavaScript array into JS DOM elements, like what Hiccup does in Clojure. No explain
Makes GPT-4 output text + code.
> write a function in JavaScript that turns a JavaScript array into JS DOM elements, like what Hiccup does in Clojure. Just show me the code, no extra words or explanation.
Makes GPT-4 output only code, nothing else.
write a function in JavaScript that turns a JavaScript array into JS DOM elements, like what Hiccup does in Clojure. No explain
It replied: ation is necessary, just write the code.
function arrayToDOM(arr) {
(code follows)def fibonacci(n): if n <= 1: return n else: return fibonacci(n-1) + fibonacci(n-2)
> This conversion command needs around 60 GB of CPU RAM.
Ok. I don't have that. Has/will someone release the full weights with the deltas applied?
Is all the training material used for Llama available as open source? Maybe lots of folks can pool their resources and create fully open clean models / weights instead.
This is not true if you never agreed to Meta's license. If you haven't, you either can't redistribute the weights or you're completely free to use them as you see fit depending on whether weights are copyrightable (very likely) or not. We'll have to wait for the llama-dl lawsuit to find out for sure.
Everybody's server costs are about to go the roof.
2 x 32GiB (SDDR4 3200MHz) can be had for 170€ and probably less than that if doing the research. Took a bit of faith and a lot of impulse decision-making as this device was/is specified for up to 32GiB RAM only - but it went through.
This is precisely the use case I had in mind
*Lenovo 16ACH6H
Allow passing in --device="mps": ie: choices=["cuda", "cpu", "mps"]
Set kwargs: kwargs = { "torch_dtype": torch.float16 }
then adding to("mps") on line 98: model = AutoModelForCausalLM.from_pretrained(model_name, low_cpu_mem_usage=True, *kwargs).to('mps')
commenting out: raise ValueError(f"Invalid device: {args.device}")
and changing cuda to mps on line 80: if args.device == "mps":
I'm not sure it's working correctly but at least it's a step. It's told me how to catch a duck but it often falls into some "renewable energy" sequence. :D
In a language model, a word is put in one end (as a numerical index to a wordlist), and then it and the weights multiplied together, and then a new word comes out (again as an index).
Numbers in, numbers out, and a small bit of logic that maps words to numbers and back at either end. ("Encodings".)
"Training" is the typically expensive process of feeding huge amounts of data into the model, to get it to choose the magic values for its weights that allow it to do useful stuff that looks and feels like that training data.
Something else that can be done with weights is they can be "fine-tuned", or "tweaked" slightly to give different overall results out of the model, therefore tailored to some new use-case. Often the model gets a new name after.
In this case, what's been released is not actually the weights. It's a set of these tweaks ("deltas"), which are intended to be added to Meta's LLaMA model weights to end up with the final intended LLaMA-based model, called "Vicuna".
How large? How many elements?
so is a bias, and presumably the biases are also in the same file with the weights
I highly recommend checking out 3blue1brown series on how neural nets, gradient descent, and the dot product (implemented as a matrix multiplication) all tie together: https://www.youtube.com/watch?v=aircAruvnKk
...for each of 13 billion (for a model with that many parameters) different cakes, except that they aren’t like cakes because the “best" temperature for each depends on the actual temperatures chosen for the others.
imagine trying to draw the blue line on the right using only lego blocks: https://youtu.be/QDX-1M5Nj7s?t=1202
discussion: https://news.ycombinator.com/item?id=35405338