AI and Open Source in 2023
magazine.sebastianraschka.com
magazine.sebastianraschka.com
I know it rustles purist feathers, but I don’t understand why we live in this pretend world that assumes that folks particularly care about respecting licenses. Consider how little success that the GNU folks have had with using the courts for any enforcement of their licenses, and that’s by stallmans own admission.
AI is itself a subversive technology, whose current versions rely on subversive training techniques. Why should we expect everyone to suddenly want to follow the rules when they read a poorly written restrictive open source license?
Re success of free licenses, linux (other than a few arguable abuses) has remained free and unencumbered thanks to GPL licensing.
It's like the difference between a bunch of shops having security cameras that they can look at the footage from, versus having every such camera connected to a wide surveillance network with facial recognition and querying abilities. Both are technically just a bunch of cameras in the same places, but not many people would argue it's the same thing.
The AI bros who celebrate this development probably think the benefit they will extract by committing this on others exceeds the negatives of others committing it on them.
Good luck.
But we also don’t know if there’s a trick up the sleeve that allows any of those models to be asked a very specific question that will result in a very specific answer.
Similarly to how map makers hide mistakes in their maps that they can point to when they want to prove that someone took their maps and presented it as their own.
Imagine using Llama for a customer support bot with a custom prompt, but there exist some phrase you don’t know that will make it say something that only Llama would say in response to that.
So sanitize user input. Don't let the question get asked the specific way the user wanted it to.
Take a page out of OpenAI's book and sanitize the output too.
Some might argue that a "pure" open-source would require the dataset and the training "recipe" as it would be needed to reproduce the training, but it would be so expensive that most people wouldn't be able to do much with it.
IMO, a release with open weights without the "source" is much better than the opposite, a release with open source and no trained weights.
And it's not like there was no progress on the open dataset front: - Together just released RedPajama V2, with enough tokens to train a very sizeable base model. - Tsinghua released UltraFeedback which allowed more people to align models using RLHF methods (like the Zephyr models from Hugging Face) - and many many others
[1] https://mistral.ai/news/announcing-mistral-7b/ [2] https://github.com/togethercomputer/RedPajama-Data
Why not both?
Like, a list of sources used and how they were harvested?
I've actually argued the opposite. https://www.marble.onl/posts/considerations_for_copyrighting... When you look at the freedoms underlying open source, you can exercise them without the training data.
Most of what disqualifies the licenses I mentioned above from being classically open source is use restrictions.
Tuned versions outperform 13b vicuña, wizard etc
https://stability.wandb.io/stability-llm/stable-lm/reports/S...
Disagree, this has long been a problem, alongside all the other familiar deceptive-but-legal false advertising out there. Fortunately the HackerNews community knows enough to call it out. Last year I posted a big list of this happening on HN: https://news.ycombinator.com/item?id=31203209
There is this ridiculous idea of AI moats and other machinations for the next big VC thing (god bless them, people have spend their energy on worse pursuits) but in a fundamental sense there is a public good type infrastructure crying out to be developed for each major linguistic domain.
Maybe such an effort would not be cutting edge enough to power the next corporate chatbot that will eliminate 99% of all jobs, but it would be a significant step up in our ability to process text.
Cited reasons: Lack of resources, lack of maintainer time and there being many new good alternatives.
Idk why nobody ever talks about it
We also have our carper.ai lab for the rl buts
We are rolling out open language models and datasets soon for a number of languages too, see our recent Japanese language models for example
Got some big plans soon, have funded it all ourself but sure other would like to help
You can try very hard to make neural network stuff a holistic social experience. There is a lot of value in that! I think it's meaningless though, a colossal waste of time.
In the objective reality we live in: We wouldn't be talking about transformers, attention, etc. if it weren't for papers that used so called "not" "open data."
It's all tainted. There's no shortcuts.
If you buy into holistic social experiences as an essential part of your chatbot or whatever, you expose yourself to being sniped in some basic way by merely one comment on the Internet. Bullshit Street is a two lane road.
Sharing research stops the second REVENUE starts flowing in. AI was always funded by VC/Research budget. When you are on a timer without a viable product, its beneficial to open source research to accelerate the time to reach a sellable product.
With GPT3.5, that point was reached, it is commercially very valuable, and OpenAI's revenue is very high even just as a consumer product, let alone all the API integrations that produce real, real value (All the outsourcing agents in India and Philippine are not more sophisticated than GPT4).
Now it makes money, time to make everything a trade secret. Google doesn't open source their search algorithm, and there's nothing wrong with that.
LLMs are not the end of the road, and the massive amounts of revenue generated by LLMs will fund the next generation of AIs to be developed. GPUs are closed source, doesn't stop them from advancing rapidly. We are in a high interest rate environment, and AI is still getting titanic amounts of investment, because its proven to generate real income, not some speculative bet anymore.
The idea being: "Theft" isn't about "you get something you don't own," it means "you deprive someone else of THEIR property."
Even if it became impossible to train AI on internet-accessible data, there's no change to the proliferation of generative AI other than keeping it entrenched and centralized in the power of a few players, and it has no impact on potentially taking jobs from artists, other than making it harder for artists to compete due to the lack of open-source alternatives.
Home weaving does still exist. Not a very big employer any more, though.
The point I was aiming for is that even if your stuff isn't ever stolen (regardless of what exactly you mean by that), you're still going to be out of a job because the automaton is good enough to out-compete you economically.
At least, that's my expectation, though I hope not for a few more years at least.
[0] yes that pun was deliberate
Technically speaking if ai would be as capable as you think it would mean that we wouldnt need said corporations since ai can provide all the services they do - otherwise it would be illogical and contradictory to claim ai can do everything under the sun.
My beef is with the oligarchy thats eroding our freedoms, privacy, and ownership. You should be too because our democracy and way of life is threatened by them. Ai is just another tool they want to monopolise.
(The personal profile/ad tracking/face tracking AI probably still could be used against people, but that doesn't seem to be part of this thread? Correct me if I'm wrong).
> Technically speaking if ai would be as capable as you think it would mean that we wouldnt need said corporations since ai can provide all the services they do - otherwise it would be illogical and contradictory to claim ai can do everything under the sun.
That requires all economic tasks to be simultaneously equally susceptible to AI. So far, not so, though this is certainly the goal.
> Ai is just another tool they want to monopolise.
Can you name a single thing that would go less badly, in the event that all the AI were made public domain?
Second, our species didn't evolve into being in 1787:
"""
The Calico Printers’ Act 1787[19] was the first statute to explicitly provide protection for designs, conferring rights enduring for 2 months upon:
“Every person who shall invent, design and print or cause to be invented, designed and printed and become the proprietors of any new and original pattern or patterns for printing linens, cottons, calicos or muslins…”
…
[19] The full title of the Calico Printers’ Act 1787 was “An Act for encouragement of the Arts of Designing and Printing Linens, Cottons, Calicos and Muslins by vesting the properties thereof in the Designers, Printers and Proprietors for a limited time”.
"""
- https://assets.publishing.service.gov.uk/media/5a7d90b2e5274...
Since the moment you put it up online for people to see and hear, they were able to move on and create something else based upon this. Most of the time unconsciously. This is how humanity works. This is the reason we're still on this planet. AI accelerates the process like any other tool we've come up with since we climbed down the trees.
You can complain and scream as much as you want, but it won't change. Even if you manage to regulate the whole western part of the internet. The rest of the world is bigger and won't sleep.
Policy, education, law, etc.
As long as there is a massive fundamental asymmetry between assembling a chip with a small amount of ROM and disassembling & reading that ROM while still making the chip usable, DRM schemes using PKI methods will become widespread and nigh unbreakable.
As correctly noted below, if the eye can see it or the ear can hear it, you have no meaningful DRM.
Of course it’s possible to process these out, same as an image watermark.