AMD Announces "Instella" Open-Source 3B Language Models
phoronix.com
phoronix.com
LLaMA 2: https://arxiv.org/pdf/2307.09288
LLaMA 3: https://arxiv.org/pdf/2407.21783
But the original poster said being glad «they even trained it on open datasets», not glad "they told you what they trained it on".
I do take personal affront when they dump models without specifying their datasets. They're just polluting the information space at that point.
There are two types of "training on copyright material"
1.) Training on material that is copyrighted and behind a paywall, but you circumvent the paywall. This is unambiguously illegal, as the material is pay to view.
2.)Training on material that is copyrighted, but free for anyone to consume. This is ambiguous in legality, but right now seems to be leaning in AI training favor - as long as the models don't verbatim share the material.
There is also another point about copyright material being ad-supported, and obviously the AI doesn't view/care about ads. There is a decent case to be made that this is illegal, but then is ad blocking actually theft?
The point is not about any «AMD's fault». It is about "why would it be great to have LLMs trained on limited (open) data".
https://github.com/allenai/OLMoE
> The Instella-3B models are licensed for academic and research purposes under a ReasearchRAIL license.
Huge mistake.
This would have been an amazing PR win for AMD if they just gave it away.
Open models attract ecosystems. It'd be a fantastic sales channel for their desktop GPU hardware if they can also build increasing support for ML with their cards and drivers.
It comes off a bit... dubious. The prohibited uses boil down to "don't use Instella to be a jerk or make porn" which I expect many people will do anyways and simply not disclose their use of Instella (which, of course, is prohibited).
This is the first RAIL license I've read, but if this is the standard, this reads less like a license and more like an unenforceable list of requests.
Edit: if I were to make an analogy, this license is a bit like if curl came with a clause that said "no web scraping or porn".
Most likely a part of strategy to dislodge Nvidia as the leading AI chip supplier, and AMD is in the position to try.
How well it will work? Well I don’t know enough details about these companies to tell.
Pre training joke for ya.
The real test will be inference latency and throughput on consumer hardware, not just the cherry-picked benchmark graphs they've shared. Anyone run comparative evals against Llama 3.2 3B or Gemma-2 on identical hardware yet?
The fully open approach (weights, hyperparams, training code) is refreshing compared to the "open weights only" trend we've been seeing. This is how you actually build a community around your tech stack.
Edge deployment is where this gets interesting - having truly open small models running locally on laptops/phones/embedded without phoning home feels like the computing paradigm we should have been pushing for all along instead of the current API-gated centralization.
In my eyes, having a completely novel and reproducible model from end to end, including its dataset is great news.
So we can't egg AMD just because they did something better in some cases and worse in others?
They released a model from end to end and shown that they can compete now. Who cares about the business applications. That can come later.
The other actors abuse the definition of Open Source, too. Not only in AI, even. So, we shall denounce others with the same force, but we don't, because of the broken window theorem.
AMD is already an underdog, so whatever they do has no merit, and worthy of booing. Do masses boo NVIDIA for their abuse of the ecosystem? Did Intel got the same treatment when they were choking everyone else unethically?
Of course not.. Because they are/were the incumbents. They had no broken windows.
I don’t care what the masses are doing.
When other companies abuse the phrase “open source” they should be called out on it.
And I’m complaining about AMD’s use of it here. Just because they’re an underdog doesn’t mean they get a pass.
But at the end of the day their released model still has restrictions on what use cases you can use the model in. If this was a piece of source code rather than an AI model it would not count as open source. And just because it’s AI model I do not think it should have a different standard.
[1]: https://huggingface.co/amd/Instella-3B-Instruct/raw/main/LIC...
https://www.licenses.ai/blog/2023/3/3/ai-pubs-rail-licenses
https://www.licenses.ai/rail-license-generator
The intent of the license is to show the techniques used in the project and to provide the results in a form that they can be used to further development of other projects but not be used themselves.
The TLDR of the license is "here's the model and the sources. use it to help make your own models but don't use our model directly"
No it isn't. It's a ResearchRAIL-MS derived license with lots of additional restrictions piled on.
> The TLDR of the license is "here's the model and the sources. use it to help make your own models but don't use our model directly"
ICYMI, it's a ResearchRAIL-MS derived license and not a ResearchRAIL-M derived license. All the onerous restrictions apply to the model as well as the source code.
You can make an MS license in their license generator by selecting ResearchRAIL and checking both the model and source buttons.
The only thing you might pick on AMDs one is "no obfuscation" clause which is not controversial at all.
(Yes, I diffed with a self-generated RAIL-MS license).
Yes.
> The terms are really broad and onerous even if one wants to use it for purely non-commercial, academic purposes.
No. It's just legalese to prevent commercial use, abuse of models and prevention of responsibility. As a person who sits in an academic research center, I see no problems at first blush.
I'm not an AMD employee, so I can't tell about their API access.
People here see student startups, I see tons of non-commercial research networks from where I sit. So the license is not absurd from where I look.
To be clear, I am happy AMD is trying to get in the game. I just dont feel this model will gain as much use as it would have had with a permissive license and I hope that AMD will switch to better licenses (like Qwen or Llama models have also done in the past).
No idea if this will work or not, but it's how I learn to use new technologies.
You could of course build this using traditional software development.
(I block ads on Android using Firefox/UBO, and on the iPad using Brave).
Still way easier to have an adblock/consent-o-matic disable it for you, but it seems like it only shows you ads if you actually agreed to them.