Some might argue that a "pure" open-source would require the dataset and the training "recipe" as it would be needed to reproduce the training, but it would be so expensive that most people wouldn't be able to do much with it.
IMO, a release with open weights without the "source" is much better than the opposite, a release with open source and no trained weights.
And it's not like there was no progress on the open dataset front: - Together just released RedPajama V2, with enough tokens to train a very sizeable base model. - Tsinghua released UltraFeedback which allowed more people to align models using RLHF methods (like the Zephyr models from Hugging Face) - and many many others
[1] https://mistral.ai/news/announcing-mistral-7b/ [2] https://github.com/togethercomputer/RedPajama-Data
Why not both?
Like, a list of sources used and how they were harvested?
I've actually argued the opposite. https://www.marble.onl/posts/considerations_for_copyrighting... When you look at the freedoms underlying open source, you can exercise them without the training data.
Most of what disqualifies the licenses I mentioned above from being classically open source is use restrictions.
I know it rustles purist feathers, but I don’t understand why we live in this pretend world that assumes that folks particularly care about respecting licenses. Consider how little success that the GNU folks have had with using the courts for any enforcement of their licenses, and that’s by stallmans own admission.
AI is itself a subversive technology, whose current versions rely on subversive training techniques. Why should we expect everyone to suddenly want to follow the rules when they read a poorly written restrictive open source license?
Re success of free licenses, linux (other than a few arguable abuses) has remained free and unencumbered thanks to GPL licensing.
It's like the difference between a bunch of shops having security cameras that they can look at the footage from, versus having every such camera connected to a wide surveillance network with facial recognition and querying abilities. Both are technically just a bunch of cameras in the same places, but not many people would argue it's the same thing.
The AI bros who celebrate this development probably think the benefit they will extract by committing this on others exceeds the negatives of others committing it on them.
Good luck.
But we also don’t know if there’s a trick up the sleeve that allows any of those models to be asked a very specific question that will result in a very specific answer.
Similarly to how map makers hide mistakes in their maps that they can point to when they want to prove that someone took their maps and presented it as their own.
Imagine using Llama for a customer support bot with a custom prompt, but there exist some phrase you don’t know that will make it say something that only Llama would say in response to that.
So sanitize user input. Don't let the question get asked the specific way the user wanted it to.
Take a page out of OpenAI's book and sanitize the output too.
Disagree, this has long been a problem, alongside all the other familiar deceptive-but-legal false advertising out there. Fortunately the HackerNews community knows enough to call it out. Last year I posted a big list of this happening on HN: https://news.ycombinator.com/item?id=31203209
Tuned versions outperform 13b vicuña, wizard etc
https://stability.wandb.io/stability-llm/stable-lm/reports/S...