What kind of machine is needed to run it?
The endgame for this AI race is obvious and when it comes to AI models, open-source ones always disrupt fully closed source AI companies.
But first, we'll see which 'AI companies' will survive the lawsuits, regulations and fierce competition for funding.
The situation here is not comparable.
Naturally, as predictions go, don't quote me on that. I was wrong before.
OpenAI's "win" wasn't even so much in the research but in the design of ChatGPT as an interface. Its own model makes the same kinds of egregious mistakes as google and FB's own LLMs. Also, OpenAI was willing to just deal with the ethical fallout of releasing it into the wild with the ability to generate authoritative sounding falsehoods.
I suspect we're going to go back to a period soon where a lot of the innovation we're seeing is around interfaces and infra to make interacting with LLMs natural and applying them to product use cases where they make sense.
Microsoft and Google are a different story, they're specifically pushing these as authoritative sources of information. If we hadn't had access to ChatGPT and the ability to learn it's ins and outs, it might have taken longer to expose so may of the flaws in the Microsoft and Google services.
That repo is basically releasing a binary (the weights) and an open source runtime to run them.
It’s absurd and it damages the open source world and its real definition.
This is FSF's freedom 1: "The freedom to study how the program works, and change it so it does your computing as you wish (freedom 1). Access to the source code is a precondition for this."
By that definition I don't think you can say this model is open source then.
You could say the inference code is open source, and that's technically true (and the only thing the repo claims TBH), but calling this open source AI model is misleading.
Open source inference code is outright uninteresting. It's a few hundred lines of glue code to run a giant binary blob which does the heavy lifting, i.e. the actual computing, i.e. the actual software, i.e. the software whose source we're actually interested in.
EDIT: Actually I just noticed they didn't even open source checkpoints+tokenizer.
Even the checkpoints are provided - for free! All you have to do is ask.
Someone at Facebook spent a ton of money to train a state of the art model, open sourced the code and even provides checkpoints free of charge, and you still complain? The level of entitlement is off the charts…
That makes it literally not open source. Can I redistribute it freely? I don't think so, I might be wrong. If not, it's not open source.
> and you still complain?
But I don't complain about Facebook releasing it in whichever terms they want. I complain that people call it open source when it clearly isn't.
> The level of entitlement is off the charts…
How is pointing that it's not open source entitlement?
The model code is literally open source. Provided and licensed.
I don’t know why they haven’t released the training code - I agree it would be nice if they did - but the important thing here is they created something valuable, open sourced it, and even provided model weights - for free. Let’s appreciate it.
Even if they did publish their training code - that’s not enough to train the model. You also need their dataset. Would you still claim “the model is not open source” because the dataset is not available?
Bottom line: the model has been open sourced. The training code hasn’t - and isn’t needed by most users of this model.
With the training code you can fine tune the model with your own dataset, which exactly fits FSF's freedom 1.
It isn't "Free Software" if it doesn't allow for the four freedoms -including the right to modify and redistribute your modifications.
As a term, "open source" is ambiguous to almost having almost no actual meaning at all. But "Free Software" has a concrete meaning and if you can't redistribute something then it is not "Free Software".
Hope that helps.
Edit to add a citation: https://fsfe.org/freesoftware/
Open Source has a clear and specific definition BTW: https://opensource.org/osd/
As I see it the weights are more similar to LLVM IR. It got "programmed" by gradient descent.
At most you can say the inference code is open source (see sibling thread).
The career risk isn't worth it especially when tech is deployed client- and network-side to detect just such exfil attempts. The average of network, security, and client management staff tend to be PEs (SREs) who can code, some have PhDs, and are the cream of what was previously organized as "corporate IT" world. So I fail to see any incentive to throw away their career and reputation by giving away IP for $0.
There's a metric s*ton of optimized hardware to generate models. And I have my doubts if Sama at OpenAI, even with 10 gigabucks from Microsoft, can sustain growth, organizational culture, and long-term investment at the scale others are bringing online with less trouble and more experience.
The future interaction will AI models will most likely be through an API because the models themselves are becoming too large to fit even on the most extreme DIY NAS solutions.
TL;DR: it's not happening.
And yet I spent last night running GPT-J-6B on my desktop CPU at 2 tokens / sec. People are finally starting to optimize these models, and there's a ton of optimization to go. We'll definitely be running these locally in the next few years. This model especially looks like an ideal candidate for CPU optimization, given the pairity with GPT3, and that it's within spitting distance of the size of models like GPT-J-6B.
e.g. Toolformer: https://arxiv.org/abs/2302.04761
which uses APIs and functions to improve GPT-J beyond GPT-3 for various tasks