BLOOM: The largest open multilingual language model
bigscience.huggingface.co
bigscience.huggingface.co
Thanks to the researchers, institutions and the French government for providing the resources to make that happen. I hope more countries on the continent follow this model of open access and funding for AI research.
Though it requires an account, unlike other models on the hub (I guess they're hoping to get new users from the hype of this model).
Bloom looks pretty exciting. It's reportedly as performant as GPT-3 (I haven't tested it enough to confirm, but what little testing I did gave okay results), but the model is completely open-source; if you can afford some cloud compute, you can just upload the model on your preferred cloud provider and just generate whatever you want from it, good or evil.
I'm expecting the coming 12 months to be pretty interesting for text generation.
git clone https://huggingface.co/bigscience/bloom
You will need Git LFS and ~330GB of free space though.It is possible to run inference locally, even without very much RAM/VRAM (though it will be very slow). Get the most recent versions of "transformers" and "accelerate" from git, and then:
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained('./bloom')
inputs = tokenizer("Hacker News is", return_tensors="pt")
model = AutoModelForCausalLM.from_pretrained('./bloom', device_map="auto", offload_folder='offload', torch_dtype=torch.bfloat16)
output = model.generate(inputs["input_ids"].to(0), min_length=30, max_length=30, do_sample=True)
print(tokenizer.decode(output[0].tolist()))
After the model is loaded you can also look at where it put each part of the model (GPU, CPU RAM, or on disk): print(model.hf_device_map)
Here's a screenshot of it in action on my machine, which has 256GB RAM and 2xA6000 GPUs:Comment 1: git-lfs is now close to saturating my gigabit ethernet connection, around 990Mb/sec
$ du -csh pytorch_model*.bin | tail -n 1
329G total
$ du -sh .git
329G .git
Not sure if there's a way to get git-lfs not to do that...I believe there are performance enhancements planned in Accelerate that will allow layers to be preloaded in parallel as other layers are executing inference, which should make things faster: https://github.com/huggingface/accelerate/issues/512#issueco...
File "/home/dek/miniconda3/lib/python3.9/site-packages/accelerate/utils/offload.py", line 25, in offload_weight
array = weight.numpy()
I verified my numpy knows nothing of bfloat16:>>> import torch >>> t=torch.arange(0,6).resize(2,3).to(dtype=torch.bfloat16) /home/dek/miniconda3/lib/python3.9/site-packages/torch/_tensor.py:586: UserWarning: non-inplace resize is deprecated warnings.warn("non-inplace resize is deprecated") >>> t tensor([[0., 1., 2.], [3., 4., 5.]], dtype=torch.bfloat16) >>> t.numpy() Traceback (most recent call last): File "<stdin>", line 1, in <module> TypeError: Got unsupported ScalarType BFloat16
Seems par for the course that I'd be able to get the world's most advanced NLP model downloaded over gigabit to my home supercomputer only to be stymied by a data type.
Then it failed again:
File "/home/dek/miniconda3/lib/python3.9/site-packages/accelerate/big_modeling.py", line 188, in dispatch_model
main_device = [d for d in device_map.values() if d not in ["cpu", "disk"] [0]
IndexError: list index out of range
Now I'm curious just how long it will take to repro this error (IE, running again, with the offload files aalready written). It's also puzzling since I set the device_map to 'auto'.Again, all par for the course and stuff I expected (having worked in HPC/ML/science for 3 decades, you get used to research codes).
https://github.com/huggingface/accelerate/commit/f13c59f91e2...
This model is larger than anything HF has done, so resource usage likely has to be constrained a bit (GPT-3 was restricted at first for similar reasons, in addition to the potential abuse issues)
It seems that the developers' position is that it's not fully open-source, and that some potentially evil or unethical uses are actually restricted:
Practically speaking, though, you can just download it on a private server and do evil and afaict nobody can stop you or know you're using BLOOM.
That's not considered open source, though.
https://opensource.org/osd-annotated
Also compare
https://web.archive.org/web/20130203112329/http://dev.hasenj...
(although IBM wasn't asking specifically about the open source nature of the license).
The asterisk says "I'm not allowed to say it's open source because some website I don't care about says it's not the right definition".
I think these big language models might get closer to something we would all call sentience once they can refer directly to their weights' previous activations for context instead of just tokens.
Why would that necessarily be insufficient for sentience? Sentience is just about the ability to "feel". Whatever "feeling" is, mechanistically, it will be some process that accepts some information and produces some output for these "feelings". It's not obvious that any kind of "memory" has to be part of the input to that process. For all we know, the transformer model is producing feelings as it computes, ie. that the next generated word is selected because it "feels right", where "feels right" simply means it's the most statistically likely word to follow the current token. Can you prove that this isn't basically what that gut feeling is for people?
We don't actually know what "feelings" are fundamentally, so most of the people dismissing LaMDA's sentience don't really understand the can of worms they just stepped in.
I bet feelings are value functions - they estimate the expected future reward from a given state. As we move from self-supervised to RL, AIs will get feelings. But they are always task dependent, the system of values is built around goals and past experience, it's not something that can be separated from the specifics of the case.
Humans come preloaded with reward signals that guide the development of feelings. We got our goals in-built - to survive, to thrive, to self replicate, and then all the sub-goals necessary to achieve the main directives - to move around and manipulate objects, to be social, to obtain food, etc.
Maybe we need to build an android baby and raise it as part of human society to get similar benefits for AI. As it stands, AI has no "skin in the game" so to speak, it doesn't optimize rewards and has nothing to lose.
Predicting the next word has its limits, a model can't design and try new experiments like human agents, a static training dataset is dead while the world is alive. Even AlphaGo was limited while doing imitation learning, but surpassed humans training from scratch by accessing a Go environment (table + opponent). The environment is the real teacher.
Still, it is not all great. OpenAI, DeepMind, Google, etc. continues to produce models that are completely closed and thus can never be the subject to analysis from the rest of the community. Models which all dwarf GPT-3 at this point. My fear is that those of us arguing for access will continue to play catch up – maybe even forever.
However, while BLOOM is more open than say FAIR’s OPT [1] – which had the gall to call itself both “open” and “democratising” despite restricting commercial usage in its “open” license – I believe there is a discussion to be had about how they (and many others in the community) use the word “open”.
[1]: https://ai.facebook.com/blog/democratizing-access-to-large-s...
While I am not a lawyer, I am somewhat familiar with licenses and have gone through the BigScience RAIL License v1.0 [2], the associated blog post [3], and some background papers and documents to better understand how Hugging Face and its community motivate the licensing.
[2]: https://huggingface.co/spaces/bigscience/license
[3]: https://bigscience.huggingface.co/blog/the-bigscience-rail-l...
From the license itself: “[T]his License aims to strike a balance between both [open and responsible AI development] in order to enable responsible open-science…” I am of the opinion that this is impossible, as the definition of “responsible” they use (see Appendix A) is directly at odds with every definition of “open” that I am aware of. Furthermore, I find phrasings such as “Although the BigScience community does not aim to impose its values on potential users of this Model, it is determined to take tangible steps towards protecting the community from inappropriate uses of the work being developed by BigScience.” confusing, as it both claims not to seek to impose its values and to seek to impose its values in the very same sentence!
In essence, I wish entities such as FAIR and Hugging Face would stop using the term “open” when they so clearly disagree with all accepted definitions of it. In the blog post related to their license [2], they explicitly state that they are incompatible with the OSI definition of open which for example the ethical source movement [4] also agrees with and thus avoids labelling themselves as “open”.
[4]: https://ethicalsource.dev
Okay, so BLOOM is not “open” as in “open source”, then what is it “open” like? Putting limitations on usage makes it run afoul of definitions of “open access”: “free availability and unrestricted use” [5]. Although I have to admit that I am less familiar with how this community would view imposing ethical considerations on readers of science, I am fairly certain that barring access to scientific information based on anything other than very clear ethical considerations (say, easily-deployable, pandemic-level virus mutations) would run afoul of a great majority of the open access movement.
[5]: http://legacy.earlham.edu/~peters/fos/overview.htm
Right, so it is not “open” as in “open source” and nor “open” as in “open access”. How about “open science”? To the best of my knowledge there is no widely accepted definition of open science, but most seem to model themselves around open access, open source, and open data. Thus falling back on OSI’s definition and to quote the Open Knowledge Foundation: “Knowledge is open if anyone is free to access, use, modify, and share it — subject, at most, to measures that preserve provenance and openness” [6]. Alternatively, their short definition makes it even clearer: “Open data and content can be freely used, modified, and shared by anyone for any purpose” [7].
[6]: http://opendefinition.org/od/2.1/en
[7]: http://opendefinition.org
In summary, I do not believe that you can argue that anything that is “ethically licensed” is “open” by any reasonable definition. This however is fine, as anyone is free to dictate how the fruit of their labour is to be used. However, even when I am being charitable, it is difficult for me not to feel that there is a desire to ride on the coattails of the positive semantics attached with the word “open” that has taken arguably more than 30 years for others to build up and I feel that it is ethically questionable to muddy the waters around a term that others have worked hard to define and build communities upon. You are “ethical”, “responsible”, or some other nice term, but not “open” – own it.
Lastly, why do I as a researcher in this area object? Especially given that I think I can agree to every single ethical point in Appendix A.
Firstly – as I have argued above – it annoys me greatly that multiple entities outside the “open” movements use the terms frivolously and I believe to their own benefit. This feels like appropriation if ever I saw it.
Secondly, I believe that despite their good intentions multiple efforts motivated by ethics are reversing the direction of openness of where science has been heading over the last twenty years. Even if only marginally so.
Thirdly (and lastly), I believe these efforts to achieve their ethical goals will prove to have little to no effect and that efforts are more strongly warranted elsewhere. Actors with the capability to cause harm will either be able to ignore the license or are likely to have already obtained these kinds of models independent of publicly released models (not to mention that cutting-edge models are developed solely by multi-national corporations without ethical quandaries). Thus, to me, taking a legalistic approach is akin to paying for indulgences and staying in our academic ivory towers. Rather, I think successful initiatives to restrict potential harm much be wider in scope akin to the Campaign to Stop Killer Robots [8] or, even better, to educate the general public about these models and work on developing countermeasures to support public trust and communication once these models inevitably become commonplace are necessary and receiving far too little attention in the community in favour of hypothetical risk and legal word play.
[8]: https://en.wikipedia.org/wiki/Campaign_to_Stop_Killer_Robots
Unrelated notes: While browsing the license I noted that it has patent clauses akin to Apache 2.0, which I found interesting. It also contains a somewhat vague requirement to keep your models in sync with the upstream: “You shall undertake reasonable efforts to use the latest version of the Model.” I assume that this is intended to be used if the model is “patched” to restrict harm. But it creates another ongoing relationship between users of the model and Hugging Face.