Parti: Pathways Autoregressive Text-to-Image Model
parti.research.google
parti.research.google
This is definitely going to become commonly available tech this decade.
> This leads such models, including Parti, to produce stereotypical representations of, for example, people described as lawyers, flight attendants, homemakers, and so on...
> Models which produce photorealistic outputs, especially of people, pose additional risks and concerns around the creation of deepfakes.
> ...the range of outputs from a model is dependent on the training data, and this may have biases toward Western imagery and further prevent models from exhibiting radically new artistic styles...
> For these reasons, we have decided not to release our Parti models, code, or data for public use without further safeguards in place.
So maybe they're partly afraid that it's too powerful, but even moreso they are afraid it's too white.
It doesn't take much imagination to see these issues being used by AI opponents as justification to ban or heavily regulate machine learning research.
But in honesty, it is probably how people react. We saw this with Pulse, GPT, and many others. The authors are clear about the limitations but people talk it up too much and others shit on it. There's also a reproducibility crisis in ML (many famous networks, like Swin[1][2][3], can't be reproduced (even worse when reviewers concentrate on benchmarks)). It isn't like many can train a model like this anyways. It gives them benefit of the doubt and maintains good publicity rather than controversial.
Of course, this is extremely bad from an academic perspective and personally I believe you should have your paper revoked if it isn't reproducible. You'd be surprised how many don't track the random seed or measure variance. We have GitHub. You should be able to write training options that get approximately the same results as the paper. Otherwise I don't trust your results.
[0] https://github.com/lucidrains/parti-pytorch
[1] https://github.com/microsoft/Swin-Transformer/issues/183
[2] https://github.com/microsoft/Swin-Transformer/issues/180
[3] https://github.com/microsoft/Swin-Transformer/issues/148
This is not true for some of the other recent headline papers: DALL-E2, PALM, Imagen etc. datasets are the primary deterrent, model details are well known.
As a mortal, there's not much to learn from these insanely big models anymore, which makes me kinda sad. Training them is prohibitively expensive, the data and code are often inaccessible, and i highly suspect that the learning rate schedules to get these to converge are also black magic-ish...
However, you are fully right that the computation costs are very high.
One thing we can learn is: It really works. It scales up and gets better. Without really doing anything special. This was kind of unexpected to most people. This is really interesting. Most people expected that there is some limit and the performance would level out. But so far this does not seem to be the case. It rather looks like you could scale it up as much as you want to get even better and better performance without any limitation.
So, what to do now with this knowledge?
Maybe we should focus the research on reducing the computation costs. E.g. by better hardware (maybe neuromorphic), or more computational efficient models.
So what? Are you just taking the piss? You are saying a literal Oracle wouldn't be impressive because the "learning rate schedules" are black magic??
Also, it's not quite I'm secret - the big labs continue to publish papers on the major advancements (at least, as far as we know), which in itself enables open source.
Lastly, academia is lagging but is making efforts to get in on making such models - see Stanford's CRFM, which I believe released some models recently.
PS I rather disagree with 'AI god' as something that is likely to come out of these efforts anytime soon personally, though that's a longer discussion.
I thought the order defines how much someone has contributed. Core contribution sounds like it should be the most, so it should be first but it is not here. Equal contribution sounds like it should come right behind each other in the order but this is also not the case here.
Browser extensions like Unpinterested tell a story.
The close-to-ideal result would have been whatever google gave me 5 years ago before they started their latest ML push with query embeddings.
Relevant discussion from the last model (imagen)
Is this a clear sign of organizational dysfunction and I should sell my shares? Or am I missing something?
Do they do object detection prior and if so, at what granularity (toes, eyes, hat, gloves, etc)? Is it at the pixel level?
What evidence do you have for claiming this motivation?
> For these reasons, we have decided not to release our Parti models, code, or data for public use without further safeguards in place.
The problem is the results align too much with a western liberal viewpoint, which is anti-thetical to the western liberal viewpoint. They would prefer an AI which has a culturally diverse output.
I believe it's more complex than that but it's undeniable there is a cadre of Twitter AI activists who would pick models like this apart if released and use the worst anecdotal examples to generate a ton of bad press over "racist AI" which is why you don't see these made public.