PaliGemma: Open-Source Multimodal Model by Google
blog.roboflow.com
blog.roboflow.com
https://ai.google.dev/gemma/terms
I was hopeful it was a change of license to be FOSS.
What's the difference between "Outputs of Gemma"/"Outputs By Gemma" (both included in "Model Derivatives") and Outputs ("not deemed Model Derivatives").
There is no small nuance to intellectualize about here if 98% of the discussed thing is the contrary of what you claim it is. (Now, one could argue it is 90%, won't help)
... and even if you don't accept the OSD, what these people are distributing isn't source code in any sense at all. Nor do their restrictions usually resemble what the average person would think of as "open".
https://openai.com/policies/terms-of-use/
> What you cannot do. You may not use our Services for any illegal, harmful, or abusive activity. For example, you may not:
> Use Output to develop models that compete with OpenAI.
Training ChatGPT on the entire corpus of human creation is "fair use", but training a model on ChatGPTs output is "illegal, harmful, or abusive activity".
it doesn't say that and terms of service violations aren't illegal per se.
either way, everybody is ignoring these restrictions and OAI is way too afraid to try to enforce it for the reasons you've mentioned
Especially amusing because, although you can argue about whether training a model is an infringment, most of the training data actually have copyrights.
ML models, on the other hand, are not works of authorship and have no copyright protection at all, period. Or at least that's the natural interpretation of copyright law as it's been applied everywhere up to this point. I suspect that the large commercial interests involved will manage to buy either favorable (but bogus) court decisions or favorable (but stupid) legislation.
> Making automated decisions in domains that affect material or individual rights or well-being (e.g., finance, legal, employment, healthcare, housing, insurance, and social welfare);
idk, vague language like this that is not court-tested is scary and will cause enterprises to think twice
That doesn't mean that only open source things should be allowed to exist, just that things which are not open source should not be called open source.
Who's definition. Mine, yours, your uncle Bobs? There is no legally defined version of the term, and the term is not trademarked/registered so at this point it is a free for all. Feel free to engage in some court cases to lock the term down at your expense.
https://en.m.wikipedia.org/wiki/History_of_free_and_open-sou....
You might be better off linking to the definition of open source, per the OSI, since 2006:
Eitherway, this is why open source zealots tend to specifically say F/LOSS (free / libre open source software), to avoid any ambiguity with people who like to claim 'open source' is ambiguous and easily conflated with 'source available'.
...That there are so many of them proves their own point, I guess!
There is no true absolute definition of anything, there is at best consensus. And yet, somehow, we are able to communicate with each-other.
There isn't really much of a catch 22 here, because there's relatively little good faith debate about the definition of open source actually. The OSI definition exists and it's the definition that is most accepted by the people who use the term. If the amount of disagreement necessary to justify this line of question is "any at all," then there's legitimately no meaningful definition of any word.
By all means, feel free to debate the definition of open source, with the knowledge that technically there is no absolute definition. I will not be participating. All this sort of pedantry serves to do is try to derail valid points. If you want a different definition, pick a different term.
I mean, if you want to argue the point, I live in the US where we have a very large stack of case law that defines the definition of trademark, in which if you violate said trademark cases in a court will occur against you. And if you act belligerent about it, the court will dispatch a guy with a gun. So I recommend asking the guy with the gun (that is the right of the state to be sole arbiter of violence).
Now, as it comes to the term open source, there is a very wide range of accepted behavior that you can get away with when using the term before a guy with a gun shows up and asks you to stop regardless of what you or the OSI says.
> Now, as it comes to the term open source, there is a very wide range of accepted behavior that you can get away with when using the term before a guy with a gun shows up and asks you to stop regardless of what you or the OSI says.
So the definition of words is defined by if people shoot you if you misuse them? How are we even communicating right now if that's the only way words have any concrete definition?
Also, I may be behind on my knowledge of U.S. law, but I don't think they enforce trademark law by firing squad.
Consensus isn't perfect, but it is legitimately the best we have. It is true that misusing the term "open source" won't result in people with guns knocking on your door, (though frankly, I'm not sure what point that proves given misusing MOST words won't do that either, including words whose definitions are ostensibly the purview of government,) but it is a term that does have the benefit of not one, but multiple organizations trying to push a reasonable definition that upholds what they feel underscores the values of open source.
Because of that, "open source" generally conveys a sense of trust to the average person. It has value. This effort is why Google and others covet using the word "open source" in press releases: the average person may not fully grok the implications of open source licensing and ideals, but they sure do get the gist. Because they've heard of Linux, or The GIMP, or Krita, or OpenOffice.
That's why this is always an important issue. Language lawyering the word open source is extremely important. I understand that VC-funded SaaS companies are upset when people use their software as intended by the open source licenses they use, but they don't get that there is no having your cake and eating it too. There is no magic "license that is open source but also somehow protects my business model". Whenever an ostensibly-open-source piece of software gets rugpulled, it dilutes the term "open source". It's a magic trick, wherein suddenly not only are the users who did nothing wrong put into a weird situation, but also all of the open source contributors got tricked into contributing to a proprietary piece of software, when they absolutely would not have if they knew that was going to happen. (That said, I really do hope this pushback eventually kills the CLA scam.)
(Note: interestingly, there was an article I saw going around that was trying to pull the fool's errand of redefining "rugpull" by trying to quantify if it "counted". This is absurd in my eyes. It's a rugpull because someone was standing on the rug you just pulled. They used and contributed to software under one license, and now it's under another. It doesn't matter if it's "fair" or "necessary". Maybe if being open source isn't so important, they can just stop releasing things as open source to begin with? But that's the rub, because "open source" has immense value for getting your foot in the door, but now when people think "open source" you're gradually training them to think "for how long?" and it will damage legitimate projects that really aren't ever going "closed".)
Since police officers won't come to your house with guns if you misuse the term open source, somebody is going to need to defend it from being diluted if they want it to actually mean something. (And while police officers with guns is pretty scary, the guarantee that I will drop an unhinged 5 paragraph rant at anyone who disagrees with my take on the term "open source" is pretty terrifying, too, I imagine.) The people who use, contribute to and value open source as a concept are the ones who have the biggest incentive to be the gatekeepers of the definition, and gatekeep we shall.
That doesn't mean there is absolutely no room for disagreement on what open source means, but there's a wide gap between good-faith debates about open source ideals and values and "who let you decide what open source means and not my Uncle?".
Some contend that "open source" is confusing to laypeople. I agree, but there is no 100% solution. Terms like "free and open source" and "free software" and "libre" all have their pros and cons. The trouble is that you can't literally consolidate the entire open source definition into a two-word phrase. Also though, it's not fair for the burden of understanding technical jargon like "open source" to fall on laypeople. They should just be able to count on open source as a positive signal of trustworthiness, and the more technical people among us can do the work of trying to defend laypeople from bullshit. This story has fallen apart a little over time (laypeople are more inundated with bullshit than ever before) but still, it is the ideal.
This means you could build something on the model today which becomes disallowed tomorrow.
These people need to be called out relentlessly until they stop misleading the public.
So much for the Altmanian infinite progress of humanity. They can't set aside their ego, their money and their power, so let's put them back in the Evil Corporate Billionaire category that such people belong to.
I don't see why anyone needs to care how they define the term.
They will now start fully leveraging their distribution advantage across products and platforms.
Hugging Face blog post similar to OP: https://huggingface.co/blog/paligemma
unless you need a 3b, I wouldn't overhype this.
From my own experimentation I have found it to really pack a punch for its weight. Another small model which has been very good has been https://github.com/vikhyat/moondream .
There's not really a super easy to use software solution yet, but a few different ones have cropped up. Right now you'll have to read papers to get the training recipes.
- https://github.com/haotian-liu/LLaVA/blob/main/scripts/finet...
- https://github.com/InternLM/xtuner/tree/main
- https://github.com/TinyLLaVA/TinyLLaVA_Factory
Is a pointer in the right direction, along with:
Would love it if I could use LLaVA, but don't want to spend the money on like 18 A100s for 24 hrs that they use for training it. A lot of the models using CC BY NC 4.0 datasets, like VILA, thats not available for commercial use unless you train the model yourself. This is the first time at least a research or company has been open with this info, they specifically say: only the pt models can be used with fine-tuning for commercial use.
If you build a smaller model, you should need much less than 18 A100s for 24 hours, though I don't disagree you'll need at least a few.
[0]: https://huggingface.co/liuhaotian/llava-v1.6-mistral-7b
The article shows a screenshot with a red overlay on the dog - how was that data returned by the model, did it return a co-ordinate polygon of some sort?
UPDATE: Figured it out using this tool, the mask looks like this: https://huggingface.co/spaces/google/paligemma
<loc0099><loc0092><loc0874><loc0926><seg014><seg009><seg123><seg126><seg004><seg074><seg092><seg112><seg000><seg021><seg099><seg015><seg096><seg043><seg012><seg019>
We're actively working on this!
Our ML team has noted that the segmentation masks in particular from PaLiGemma are a bit tedious and unintuitive to decode. We should be pushing out (more!) open source software that uses this model in the coming days.
Look forward to an easy way to fine-tune PaLiGemma and broader support for its task types in `inference`, the package used in the blog post.
the blog post details it but essentially to convert from PaliGemma tokens to bbox:
y0 / 1024 * h
x0 / 1024 * w
y1 / 1024 * h
x1 / 1024 * w
have not played with segmentation yet.
I ran example 'segment cat' from the example provided in the PaliGemma demo and it responded this <loc0055><loc0115><loc1023><loc1023><seg063><seg108><seg045><seg028><seg056><seg052><seg114><seg005><seg042><seg023><seg084><seg064><seg086><seg077><seg090><seg054>
no documentation explanation on how to interpret segment token,
then you will need to pay a lot more for actual experts to tell you why benchmark Z is bullshit and model Y2 is actually better for the task you're actually trying to do and btw would you like to develop your own because that's a moat.
Or you get that for free here on HN.
Also, I imagine there isnt some oobabooga/automatic1111 for this?
It blurs what is realistically useful.
Also 'GPUs' No. This is a major red flag man. They do not. The marketers told you this, and you believed them.
The iPhone does have a GPU, and it can be used to accelerate AI workloads.
This isn't theoretical. You can see this for yourself by installing https://apps.apple.com/us/app/mlc-chat/id6448482937 and turning off your wifi and running the quite capable Mistral 7B Instruct LLM on your phone.
I've even used it to answer simple questions while I was offline and only had my phone with me.
What am I missing here?
Buddy I'm cranking out 4k.
Also LOL at the GPU thing. Its insane how well Apple marketing worked.
Most people don't know that's possible yet.
Why do keep indicating that iPhones don't have a GPU?
- single interconnected neural network (LLM attention layers break this, autoencoders complicate this)
- single training pass (LLMs have multiple passes, GANs have a single but produce multiple models)
LLMs have multiple passes? wdym?
Not to mention it's built to be fine tuned and commercially permissive!
"In average accuracy, we saw 85.84%, beating all other OCR models except for Anthropic’s Claude 3 Opus."