Sharing new research, models, and datasets from Meta FAIR
ai.meta.com
ai.meta.com
There have been a few attempts at prompt-driven editing (I think it was called Instruct-something) and there's a whole field of work on stuff like animation transfer (megaportraits and emo come to mind for close-up, and there are a few things that do broad motion out there) but it's hard to suggest something without knowing your use case, and if you're looking for general purpose the research seems to just not be there yet. Like that might be the kind of thing that needs a "world model"
I think probably controlnet + some kind of LoRA embedding for your specific character is going to be the closest you can get right now, but it's definitely some extra work and not quite what you want
From what I understand, they are essentially training it to have some form of representation of the context as a whole that is then used to generate the next n tokens, I feel like this is a nice next step towards "smarter" models. I wonder if a similar thing can be done for the inputs.
It's a shame they didn't compare it to llama3 since they had both a 6.7B and a 13B multi-token model. From what I could gather, the intruction-tuned llama3 is much better on HumanEval for example.
Edit, it was a month ago: https://twitter.com/ylecun/status/1793181068943639014
Cool, it only took about two years from wire spread public releases to SV lockdown of the powerful tool and then throw the scraps to the peasants.
The overloads of social ethics have already decided what you cannot have. Just like the elders of the internet intended…
Their image tokenizer creates image encodings in a discrete space of dimension 8192; each image consists of 1024 discrete tokens. The 8192 vector size matches their language tokenizer (discrete and also size 8192). Then they just mash those tokens together, which leads to downstream problems in training: because the two token spaces have different entropies, stable training in later steps is difficult, and they have to do a bunch of regularizations to make sure that gradients don't start exploding.
Interesting stuff!
The reason for this exclusion just can't be a banned criteria such as race.
(I.e. kicking someone out after they've stirred up controversy)
I'm not saying that this happened to you, I'm just addressing the point you made, that anything "open" can't or shouldn't be exclusionary.
As for why? I have a theory: Meta is not in a position to capitalize upon the model itself. Yes, they can use it internally, and maybe their competitors can copy it to - but there are no real competitors to Facebook or Instagram that can benefit from it enough to make it a differentiating facet.
Thus, releasing stuff for open source does two things:
1) Make them more attractive to research talent (Apple famously recently started to publish research because their traditional secrecy was causing issues with hiring top talent) and...
2) Continues to undermine the ability to make $$$ off of model alone, driving it towards being a commodity rather than the long term profit engine for other companies.
This is not true. Meta Fair has been built on openness from day 1. We published many papers and open-source d many repositories to reproduce the work
For example:
Faster R-CNN - state of the art image segmentation, released in 2017.
FastText - text embedding models, 2016.
FAISS - vector DB, 2018.
https://github.com/orgs/facebookresearch/repositories has over 1,000 repos.
Per your links, it's clear that FAIR does have a good history of open source work.
As I frequently point out, Facebook are free to decide on their own culture as they see fit and I am not entitled to their work. But it saddens me that they believe that compromising on their initial ideals is the way forward, rather than sticking to them through thick and thin. This, ultimately, makes it more and more difficult for me as an academic that believe in these ideals to work with them.
You can hardly call that "leak" when they basically were sending the weights to thousands of people who applied for access. It is not that they kept them secret.
2/ Prevent OpenAI from cornering the future $$$ market. Unfortunately, Google search is hit as well, but it is more due to the generational shift.
3/ Attract the best AI researchers. A product is a good as its core set of people (often just a few).
Also, maybe they need to improve their brand. Hoarding data for over a decade, maybe bringing something back now.
They were always present at ACL with decent open research as far as I have been studying/working in NLP (2014).
It's part of a strategy to attract top talent in the field. If you want top researchers you have to let them publish, which in turn hones a reputation of solid research, attracting more talent.
This. Although, it turns out, if you pay them well enough (like OpenAI), they'll forego publishing. If you can make enough to retire by 40, why work at places that pay less?
ofc, not all researchers think that way, and there are those who are in it for the science, not just money.
They do not want their users to go to ClosedAI or similar and communicate with an Artificial Stupidity instead talking to each other on Facebook.
So it is in their interest to undermine the market for Artificial Stupidities by releasing the models for free.
Just wait until they have a fully intelligent automated pipeline for lifestyle ingestion (e.g. pervasive analysis of all communication and AV) into AI management (Timeline / Memories) backposted into web 2.0 feeds like Facebook.
Google did the same thing when they released android for free.
[1] https://www.harperacademic.com/book/9780062896322/the-busine...
thanks to meta i have been creating instrumentals of some of my favorite songs with vocals https://github.com/facebookresearch/demucs