Meta Open-Sources Computer Vision Foundation Model DINOv2
infoq.com
infoq.com
"I think that there's an important distinction between the products we offer and a lot of the technical infrastructure, especially the software that we -- that we write to support that. And historically, whether it's the Open Compute project that we've done or just open sourcing a lot of the infrastructure that we've built, we've historically open sourced a lot of that infrastructure, even though the products themselves are obviously were not -- we haven’t open sourced the code for our core products or anything like that.
And the reason why I think why we do this is that unlike some of the other companies in the space, we're not selling a cloud computing service where we try to keep the different software infrastructure that we're building proprietary. For us, it's way better if the industry standardizes on the basic tools that we're using and therefore we can benefit from the improvements that others make and others’ use of those tools can, in some cases like Open Compute, drive down the costs of those things which make our business more efficient too.
So I think to some degree we're just playing a different game on the infrastructure than companies like Google or Microsoft or Amazon, and that creates different incentives for us. So overall, I think that that's going to lead us to do more work in terms of open sourcing, some of the lower level models and tools.
But of course, a lot of the product work itself is going to be specific and integrated with the things that we do. So it's not that everything we do is going to be open. Obviously, a bunch of this needs to be developed in a way that creates unique value for our products, but I think in terms of the basic models, I would expect us to be pushing and helping to build out an open ecosystem here, which I think is something that's going to be important."
I have always had the impression that maintaining an open source project is way more work than you get back from "the community" of users. Is this not true? Are for instance the internal facebook react users benefiting a huge amount from what outside contributes have built on top of react?
I think an unspoken dimension is that kneecaping the other big tech companies' entrenchments and denying them a market is always good for them - esp when as they point out, it doesn't actually hurt any of their own business interests. Other faang are always a future threat. Hurting them is always a good business move
FB's plan is to F everyone else (MAAG) by making sure they can't make billions off tech that FB have sitting on the shelf, yet is extremely expensive for a true startup competitor to get in on.
"Paradoxically, the one clear winner in all of this is Meta. Because the leaked model was theirs, they have effectively garnered an entire planet's worth of free labor. Since most open source innovation is happening on top of their architecture, there is nothing stopping them from directly incorporating it into their products.
The value of owning the ecosystem cannot be overstated. Google itself has successfully used this paradigm in its open source offerings, like Chrome and Android. By owning the platform where innovation happens, Google cements itself as a thought leader and direction-setter, earning the ability to shape the narrative on ideas that are larger than itself."
https://www.semianalysis.com/p/google-we-have-no-moat-and-ne...
Except neither this model nor several of their recently-lauded “open” releases are open source; they are CC-BY-NC 4.0, aka, you are free to tinker and share, but not to use the work or derivatives for commercial purposes. Any community effort the Meta’s hobbyist-source license attracts is work that isn’t enabling commercial competition, unlike actual open source systems like Suno’s Bark (MIT) or even use-restricted-but-not-non-commercial shared source licenses like Stable Diffusion’s CreativeML Open RAIL-M.
So what?
Sure, maybe the Googles of the world aren't building on top of meta's products, but I can tell you that a lot of startups are.
Does it make these startups vulnerable, to long term future legal action? Sure, but nobody is thinking that far ahead. What people are thinking about is how to get users and show off flashy demos to investors.
Instead, people are just pushing out products, breaking meta's licenses, and not telling people about it, while they attempt to get traction.
Strict licensing, without enforcement, is not worth the paper that the contract is written on.
So yes, it is still beneficial that the code is released, even with a bad license.
So, that's a reason that might wish to release a non-open-source model with this particular license, and one that provides an alternative to the “Meta is doing this because they stand to benefit from open source models taking off”, specifically, “Meta is doing this because it stands to benefit from drawing energy away from open source models into ones that cannot legally be used to commercially compete”.
> Does it make these startups vulnerable, to long term future legal action? Sure, but nobody is thinking that far ahead.
Well, the startups may not be, but Meta maybe is, and its acquiring a zero-cost, upside-only investment in every startup doing that. “Unjust enrichment”.
Crying about it doesn't change it's effectiveness.
And really, I don't think meta cares either.
They likely are releasing this stuff, with a strict license, just so they don't have any liability, or bad publicity.
But, they likely are happy that everyone is using their stuff.
A better example would be Uber, though, a company which is now massively success despite early legal problems
Also, once again, Facebook likely wants everyone to be using their product and is unlikely to go after people.
Sure. And, there's an argument that the license only applies to the code because model weights aren’t subject to copyright anyway. And available-under-any-license is a lot better than OpenAI’s current stance as far as enabling anyone else, since they’ve gone completely closed to the point where even their papers on their models are more PR than reproducible science. There's a continuum from secret sauce to “do what thou wilt”, and I am not a zealot arguing anything not Open Source must be rejected as not a positive step.
In a way, it takes away their competitors edge while racing to the bottom to compete with open source. At the same time, they establish themselves as experts and keep attracting great talent that wants to publish their work openly. And it benefits all of us, so good marketing amongst developers too.
EDIT: On reflection, you can probably extend the content creation argument to say that noncommercial tools enabling that without enabling commercial competition, to the extent that some of the models will be integrated into Meta products, is the best of all worlds for Meta, so the basic argument works even without open source in the strict sense.
Facebook doesn't want the models to be the money making bit, because they aren't a licensing/subscription service. They are an ads and soon hardware-platform company. They want those bits to be what people pay for. Not the models.
All these models are licensed under a non-commercial license. So their competitors don't gain a real advantage.
Other than OpenAI (who are remarkably tight lipped), ML researchers are pretty chatty in both their papers and watercooler hangouts. So, the information is going to get out either way. Might as well get ahead of it, and look like the good guy in the process.
We've moved on from ImageNet-style tests "Choose the most appropriate label for this image from 200 possible labels" to much more advanced "Reasoning" tests[0]. PaLI[1] is potentially the SoTA here but BeIT-3[2] may be better example for my thesis. Notice that BeIT-3 is trained on not just images, but also trained in natural language. It outperforms purely image-trained models on even pure-image tasks like Object Detection and Semantic Segmentation.
Take a look at the major benchmarks for Segmentation (ade20k) [4]: DINOv2, 11th place. BEiT-3, 4th place. Yes, BEiT-3 has 72% more parameters but it's also basically an entire LLM. Even GPT-4 is a multi-modal model, and actually accepts images as prompt inputs, OpenAI just doesn't expose that ability.
More importantly, the new multi-modal models can understand human questioning like "What type of flowers are in the blue buckets of this image?" and respond intelligently, in English/whatever.
DINOv2 was trained with techniques borrowed from LLM training methods, but is not trained for natural language.
0: https://paperswithcode.com/area/reasoning
1: https://arxiv.org/pdf/2209.06794v2.pdf
2: https://paperswithcode.com/paper/image-as-a-foreign-language...
3: http://www.incompleteideas.net/IncIdeas/BitterLesson.html
4: https://paperswithcode.com/sota/semantic-segmentation-on-ade...
1) Positive PR to compensate, e.g for privacy law violation fines (see yesterdays news cycle)
2) Identifying and nurturing the next generation of hires or acquihires in this very technical area
3) Depriving the competition from alternative business models that might threaten the adtech walled garden they exclusively rely on
4) Collecting ideas about how these models can be improved or used creatively (related to 2)
5) Keeping some key employees happy with intangible rewards (name recognition, career prospects)
When you spend billions on development this type of strategic leakage is just a few marked droplets into a digital ocean
This opens up an opportunity for their competitors to eat into their moat because OpenAI is treading water/downgrading their product, chasing scale. Meta is leveraging this opening to flood the field with amazing open source tools, all of which compete with OpenAI offerings, knowing that the open source community will run with them and further erode OpenAI’s moat.
Meta is cornered. And it needs to figure its way through its current positioning.
They have some deep tech, and they have deep pools of historic global sentiment, which, if one were to train GPT4 on...
So, I can only surmise nefarious actions shall (are?) be afoot.
https://github.com/facebookresearch/dinov2/blob/main/LICENSE
The license seems to only apply to "sharing" the code. Probably because it addresses copyright issues.
But what about commercial use of the code and weights?
Can one incorporate those into a server side system which lets users pay for using it?
You may consider any output of the model derivative work and then it would apply. But that goes against the understanding that outputs of ML models are copyrightable.
If however, someone runs the model for non-commercial purpose and you take the the output of the model and commercialize it...
Copyright law is broken by now IMHO.
hard to use the model without
reproducing (copying) it
Can you link to a source that backs up this line of thought? It is not how I understand copyright law. AFAIK it is not about you making a copy for yourself. It is about person A copying content person B has the copyright to and then giving it to person C. That's why the license talks about "sharing".The way I understand it, what A does with their copy of the content does not fall under the copyright law.
I don't think this is what copyright is about. As far as I know, it is about sharing a copy with a third party. But I am happy to be corrected if someone has a link to a reputable source that says otherwise.
[0] https://itlaw.fandom.com/wiki/Temporary_copies
I remain sceptical about the idea that you can't. Look at the license in question:
https://github.com/facebookresearch/dinov2/blob/main/LICENSE
Not a single time it talks about the act of copying and/or using the software. It constantly talks about sharing it.
The consumer.org.nz page you linked to also mainly focusses on "selling" the data itself and explicitely states that you can make copies for yourself:
Can I download mp3s from the internet
and copy them to CD?
Sure – as long as they were obtained legitimately
– either purchased from an online store or from a
legal free download. As we said before, if you own
the music legally then you can make a copy.
This seems to underline my thinking that copyright is mostly about sharing with 3rd parties. Not about the act of duplicating bytes in a technical manner or about what you do with your copy of the bytes.The license talks about "reproducing", which is copying.
You seem to be basically saying that you think Meta's lawyers and Creative Commons's lawyers don't know how to do their jobs (and can be trivially out-thought in a couple of minutes by laypeople), which seems very unlikely.
the Licensor hereby grants You ... to ...
a. reproduce and Share the Licensed Material,
in whole or in part, for NonCommercial purposes
b. produce, reproduce, and Share Adapted Material
for NonCommercial purposes
As you can see, it always talks about sharing the content which can only be done for NonCommercial reasons. If it wouldn't only apply to sharing but to any type of usage, why would sharing be mentioned in every paragraph?Also, in my experience, a lawyer would use "or Share" and not "and Share" if they wanted to express that the act of "reproducing" alone (whatever that means) is enough to fall under the license, even if no sharing would occur.
So my feeling is still that the license deals with sharing. Not with act of using.
The GPLv3 is certainly an open source license, sure.
Better in blanket terms is...not the point I’m making, and I am not arguing that, in this space, all open source licenses are categorically better than non-open source licenses.
But, I do think its important to describe licenses accurately and understand the implications of particular licenses.
> However to push back a bit, I would say that at least the license allows researchers to do the important work that needs to be done.
Yes, as far as sharing research, this is worlds better than OpenAI. And its worth noting that while the usage restrictions aren’t as competition-restricting, the most widely touted successful “open source” model (Stable Diffusion) is also not open source strictly (the license has usage restrictions) though there are some notable truly-open-source models.
> 6. No Discrimination Against Fields of Endeavor
> The license must not restrict anyone from making use of the program in a specific field of endeavor. For example, it may not restrict the program from being used in a business, or from being used for genetic research.
For transparency and user freedom, CC-BY-NC 4.0 is still better than the proprietary license of OpenAI's closed source models.
Exactly. I'm for using definitions of words as they are commonly understood. The other definition relies on arbitrary conditions totally unrelated to the English meanings of the words being used.
No you are not. Open source in the context of computing has a definition that is commonly understood. Why then are you using definition of that word that is not commonly understood?
Licenses adhering to the OSD generally ensure open viewing, open use, open modification and open distribution. Stripping that back to just viewing removes a large part of openness that "open source" has been built upon.
For folks in the know: I often see segmentation models on video frames producing patchy results (see the DinoV2 video of the running dog, the body gets black patches randomly, so the segmentation fails for certain frames). What methods are folks using to deal with this - standard fine-tuning, or is there a way to "force" the area to be cleanly segmented (ie, add a bounding box around the class to supplement the data)?
And is it something that can be implemented in foundation models, or are we always going to have patchy zero-shot results like this on video files?