I've had tons of fun implementing LLaMA, learning and playing around with variations like Vicuna. I learned a lot and probably wouldn't have got so interested in this space if the leak didn't happen.
I've had tons of fun implementing LLaMA, learning and playing around with variations like Vicuna. I learned a lot and probably wouldn't have got so interested in this space if the leak didn't happen.
If it was a deliberate leak, it was a good idea.
So either, it is very hypocrite of them to apply DCMA while the model itself is illegal. Or, they are trying to somewhat stop spreading as they know it is illegal.
Anyways, since the training code and data sources are opensource, you 'could' have trained it yourself. But even then, you are still at risk for the pirated books part.
This gives the company plausible deniability while still allowing ~unrestricted growth.
Persistent storage (in violation of TOS) and illicit use of Facebook users’ personal data was available to app developers for a long time.
It encouraged development of viral applications while throwing off massive value to those willing to break the published rules.
This resulted in outsized and unexpected repercussions though, including the Cambridge Analytica scandal.
People should be wary of the development as much as they are enthused. The power is immense and potential for abuse far from understood.
For example, WhatsApp’s greyhat use of smartphone address book.
The US government also has a stake in unbridled growth seems, in general, to give a pass to business exploring new terrain.
That was ironically Bill Gates
https://www.latimes.com/archives/la-xpm-2006-apr-09-fi-micro...
You might see hackers, employees, or contractors leaking models more frequently.
And since models are distilled functionality (no microservices and databases to deploy), they're much easier to run than a constellation of cloud infrastructure.
They are generally trade secrets now, which is what actually protects them. Leaks of trade secrets are serious business regardless of the IP status of the work otherwise.
We won't know until this hits the courts.
I don’t think “Public domain” means what you think it means.
IANAL but, I think, as far as US law goes, they have the right conclusion for the wrong reasons. Unsupervised training is an automated process, and the US Copyright Office has said [0] that the product of automated processes can't be copyrighted. While that statement was focused on the output of running an AI model, not the output of its training process (the parameters), I can't see how – for a model produced by unsupervised training – the conclusion would be any different.
This is probably not the case in many non-US jurisdictions, such as the EU, UK, Australia, etc – all of which have far weaker standards for copyrightability than the US does. It may not apply for supervised training – the supervision may be sufficient human input for copyrightability even in the US. It may not apply for AI models trained from copyrighted datasets, where the copyright owner of the dataset is claiming ownership of the model – that is not the case for OpenAI/Google/Meta/etc, who are all using training datasets predominantly copyright by third parties, but maybe Getty Images will build their own Stable Diffusion-style AI based on their image library, and that might give them a way of copyrighting their model which OpenAI/Google/Meta/etc lack.
It is always possible that US Congress will amend the law to make AI parameters copyrightable, or introduce some sui generis non-copyright legal protection for them, like the semiconductor mask work rights which were legislated in response to court rulings that semiconductor masks could not be copyrighted. I think the odds are reasonably high they will in fact do that sooner or later, but nobody knows for certain how things will pan out.
[0] https://www.federalregister.gov/documents/2023/03/16/2023-05...
That output could still be covered by copyright: In the case where the input is covered by copyright, the product/output may be considered a derived work, in which case the output is still covered by the same copyright the input was. Your argument just explains why the output will not gain any additional copyright coverage.
The training code is Apache 2.0 licensed so it can be copied and modified freely, including for commercial purpoes. https://github.com/facebookresearch/llama
But AFAIK this is just the first step to get initial weights and later you need much more work to fine-tune this to get useful results from the model.
I think this step could be seen as contaminating weights with copyrighted content.
Something like chrome is copyrighted but chromium is not
I'm not a lawyer, so I'm not that well informed how official definitions match here, but what's I'm trying to say it that I wouldn't be surprised if this would go either way
But what's it gonna do in the hands of your parents or kids.. when it gets thing wrong, its could have way worst impact if it's intergrated in government, health care, finance etc..