If we use recent history as an example, OpenAI announced DALL-E in Jan 5, 2021 [3], announced v2 and a waitlist for public use in July 20, 2022, and Stable Diffusion shipped an open source model on August 22, 2022 [4] using ~$600K of compute (at retail prices on AWS) [5].
I don't see how it's likely that any company can acquire a durable technology moat here. There are scale barriers to entry, but even VC sized funding can overcome that.
[1] https://twitter.com/goodside/status/1611556749726605312
[2] https://humanloop.com/blog/stability-ai-partnership
[3] https://en.wikipedia.org/wiki/DALL-E
[4] https://en.wikipedia.org/wiki/Stable_Diffusion
[5] https://twitter.com/emostaque/status/1563870674111832066?lan...
About this license
The Responsible AI License allows users to take advantage of the model in a wide range of settings (including free use and redistribution) as long as they respect the specific use case restrictions outlined, which correspond to model applications the licensor deems ill-suited for the model or are likely to cause harm.
[1] https://huggingface.co/CompVis/stable-diffusion-v-1-4-origin...
This isn't a very minor point, as this was an explicit discussion and is also OSI's translation of Debian's translation of Richard Stallman's "freedom 0".
That is, it's an important, and explicit, tradition/consensus in FOSS that users aren't restricted in the purposes for which they may use the software.
Ok, so a major restriction on what you can do with the software.
But those definitions are clear that the "right to run the program for any purpose" must not be restricted by copyright licensing terms, and that copyright licensing "must not restrict anyone from making use of the program in a specific field of endeavor". Neither of those are infringed by restrictions on further distribution. (In fact, even freeware licenses that prohibited redistribution entirely could be compatible with this specific rule.)
You might say that it was surprising or hypocritical not to have a corresponding freedom related to redistribution, which would then preclude copyleft licensing. The BSD projects have tended to act as though they recognized this additional rule (that it's important to allow sublicensing and not to attach the same conditions to derived works, including allowing the possibility that end users of derived works will get fewer rights). But even in this case, nobody has suggested that it was "free" or "open" to directly limit the purposes for which end users could run a program.
The non-finetuned BLOOM does not appear favorably (in English) compared to GLM or OPT, which both have published weights: https://crfm.stanford.edu/helm/v0.1.0/?group=mmlu and Flan-T5 is above OPT-IML: https://arxiv.org/pdf/2212.12017.pdf
> Is the future going to be controlled by big corporations who own the models themselves?
On this subject, there is an effort stemming from BigScience to build an open, distributed inference network, so that people that don’t have enough GPUs at home can contribute theirs and get text generation at one word per second: https://github.com/bigscience-workshop/petals#how-does-it-wo...
In practice these models are typically run using top-tier A100 GPUs, which apparently is the cheapest thing you can do at scale: https://forum.effectivealtruism.org/posts/foptmf8C25TzJuit6/.... It looks like you can get away with just $10/hour, but I'm not sure I believe it. In one hour you can roughly generate 6 million English words this way, that's quite cheap.
But if you want to own the full hardware, then it's quite more expensive. You need 8 of those A100 GPUs, which come at $32k a piece, so you're in the ballpark of > $300k to build the server you need. Then there's of course running costs, these GPUs burn 250W a piece, plus the rest of the server we're at about 3kW power. That's not much, maybe $0.50/hr, plus maybe another $1/hr to cool the room it's in, depending on where it is (and the season, I guess in winter a fan might suffice, it's about as powerful as a couple small electric heaters). So with an upfront expense of > $300k, you're maybe down from $10/hr to $1.5/hr, saving something like $8.5/hr, which is $6k / month (minus the rent of whatever place you put the server in).
All in all, it's definitely feasible for a small start up as well, but not very much for an individual.
It is not the size of the model or the text it was trained on that makes ChatGPT so performant. It is the additional human assisted training to make it respond well to instructions. Open source versions of that are just starting to see the light of day [2].
” This repository has gone viral without my permission. Next time, if you are promoting my unfinished repositories (notice the work in progress flag) for twitter engagement or eyeballs, at least (1) do your research or (2) be totally transparent with your readers about the capacity of the repository without resorting to clickbait. (1) I was not the first, CarperAI had been working on RLHF months before, link below. (2) There is no trained model. This is just the ship and overall map. We still need millions of dollars of compute + data to sail to the correct point in high dimensional parameter space. Even then, you need professional sailors (like Robin Rombach of Stable Diffusion fame) to actually guide the ship through turbulent times to that point.”
[p.s. ^ was just fyi/heads-up + https://github.com/CarperAI]
That's a bizarre thing to write for a public repository
A free for all will still result in a dizzying paradigm shift though, but there is no alternative. Like Guttenberg but exponentially faster - locality will become central, "corny pop culture hacker dungeons" / hackspaces could become important as no one will know whats real and only locally controlled compute and algo power is to be trusted.
The computing bill around large language models is extraordinary (hence MSFT Azure being a strategic partner).
Among others, they generally dont think for themselves. The picture could change significantly if the various intermediaries, consultants etc who live off these ecosystems found ways to make open source profitable for them
Huggingface is one of the most beloved companies in the world right now too. Lots of interest in them within open source developer communities.
I'm likely to do less NLP research going forward and more CV research because I can't locally run most LLMs but I sure as shit can run most of the diffusion models at home.
It's a sad situation.