More than 80 AI models from Qualcomm
huggingface.co
huggingface.co
Calling freely downloadable weights "Open Source" quite diminishes the term. It's laudable, so I'd hate to discourage it. But it's not Open Source.
What constitutes Open Source is not vague, btw. it's well defined by the OSI [1].
Let's call it open weights or freely downloadable weights or something.
EDIT: I was mistaken - they are down sampled/right-sized w half precision:
BTW props to Qualcomm but these are "just" quantized versions of existing models? Useful, yes, but maybe not that novel.
Network weights and biases plus architecture models are things that can be directly build upon, just like graphics are, and I would count a photo licensed under, say, MIT, as "open source" even if the JPEG codec on the camera which took the photo was not.
2. How are these two licenses compatible with https://opensource.org/osd part 2. Source Code?
a. https://opensource.org/license/unicode-license-v3 - separately lists data files
b. https://opensource.org/license/nasa1-3-php - "Notwithstanding any provisions contained herein, Recipient is hereby put on notice that export of any goods or technical data from the United States may require some form of export license from the U.S. Government. Failure to obtain necessary export licenses may result in criminal liability under U.S. laws. Government Agency neither represents that a license shall not be required nor that, if required, it shall be issued. Nothing granted herein provides any such export license."
Large models don’t work like that.
The only practical way to modify (finetune, lora, merge) a model is using its binary form. A source dataset may be interesting, but it’s non-modifiable and non-reproducible in practice due to training costs. Rebuild process is usually non-deterministic, so “verify” is basically not an option.
So technically true, practically complicated. Open weights and open dataset would be better terms.
Some trained weights doesn't seem to qualify, to me.
Well, most of us. I'm sure there's at least one person here who can afford to burn 25 million USD on compute just for fun.
IANAL, but at most, publishing weights with a license may amount to little more than a ToS agreement, allowing distributors a bit more leeway in managing their legal/commercial relationship with recipients of said models. In other words, breaking the terms laid out in a text file entitled "LICENSE.txt" and distributed alongside a set of model weights may constitute a breach of contract, but it is in no way a copyright violation.
[1] https://fingfx.thomsonreuters.com/gfx/legaldocs/klpygnkyrpg/...
I'm pretty surprised to see objections to my claims here, I kinda thought I was just stating the obvious.
> Who is this organisation that they get to mandate the definition of the english language?
> What authority do they have to define the term “open source”?
This term has meaning and while its meaning started with an group of people forming an organization saying "this is what this means" its meaning doesn't derive from OSI (nor some farcical aquatic ceremony). Its meaning comes from its popular use in language. For example, Wikipedia does describe this term and how it came to be used [1].
IMO it would be unclear and confusing to use the same term Open Source to describe both what it has historically described and how model weights like these are distributed. The term "Open Source" itself was coined to disambiguate merely "open" source from "free-as-in-freedom" source.
Are you saying they've just repacked other existing models under their own banner but haven't opened sourced some other component?
[0] https://huggingface.co/qualcomm/FFNet-40S
> The program must include source code, and must allow distribution in source code as well as compiled form.
The compute graph with trained weight is very much a compiled form of the model. The source code would include everything needed to train that model and reproduce it.
EDIT: wasn't fast enough
You know these models are trained on internet scrape which contains copyrighted content, so the dataset can't be open sourced. It's either this or bad models.
I'm not saying "opening models is bad", it's good. However imo it would be nice to have a semantic way to differentiate between those two
They don't get to claim that it's open source just because it would be too hard to actually open source.
Can you reproduce these models? if not then it's probably not open source. With a model the closest analog seems to be the training data. Is that all published?
What authority do they have to define the term “open source”?
This post[0] dives into who coined the term (spoiler: it predates OSI by a long, long time), but it’s reasonable that OSI popularized it alongside their specific definition.
[0] https://lunduke.substack.com/p/who-really-coined-the-term-op...
IMHO, to the extent that their goal was to find a less confusing term than “free”, I’d say they’ve failed.
edit: I'm confused how this very basic claim got downvoted.
Who is this organisation that they get to mandate the definition of the english language?
Open Source is more like a designation. It is an agreed upon set of requirements that, if you change a requirement, it is something else. This is important.
Some things have legally protected designations such as 'ice cream'. Ice Cream has specific meaning in industry and even a grading system. If someone wants to make a cheaper product than the lowest grade of ice cream, they can't call it ice cream, they have to call it something like: frozen dairy dessert.
This makes it easy for people understand what they are actually getting and paying for.
I wouldn't get indignant about mandating english language definitions. I would be indignant that ai companies are not fulfilling the requirements to call it open source and are providing a cheaper product than the abilities that an actual open source model would provide.
I also encourage you take a look at our GitHub repository: https://github.com/quic/ai-hub-models
If you have questions or feature requests, you can reach out to us on Slack (https://join.slack.com/t/qualcomm-ai-hub/shared_invite/zt-2d...) or file an issue on GitHub / Huggingface. We are pretty responsive!
https://www.jnmjournal.org/journal/view.html?doi=10.5056/jnm...
Systemic review: the pathogenesis and pharmacological treatment of hiccups
Looking forward to an AI model that can cleanly remove backgrounds and jaggies from an image.
Whisper small and base are coming in the next release as well, so look out for that in the next week or two.