Implementation of Imagen, Google's text-to-image neural network, in PyTorch
github.com
github.com
As a thank you to him -- he also does for work for commission/etc, check hus GitHub page for more info. I'm not fiscally or currently otherwise directly linked to him too closely, I've just hung around a while and think he deserves far more credit than he gets. This is literally the smallest piece of the pie of what the man does across several subdisciplines, send him a thank-you please, if possible!
Why don't they just keep their findings to themselfes and build products on top of them?
Public companies can't do stuff just for the fun of it, right? So there must be some commercial reasoning behind it?
It's less cynical, more incentive alignment.
- being open is kind of just how things in ML generally work right now, it's in stark contrast to things like chemistry or physics where paywalls are pretty common
- it's a matter of clout, ML is moving ridiculously quickly, with work from just 5 years ago being considered outdated in terms of capability, if you don't publish, someone else will and they'll get the credit. This likely also matters for the researchers since they get credit too. In a sense this is just publish or perish culture from academia.
- it's also somewhat about hiring, which is related to the clout. By putting out this kind of research, they're attracting talented engineers to consider working for them. This of course is pretty relevant to the rest of their business, especially given how heavily Google leans on AI to handle moderation.
2. Deploying models in a cost effective way is hard
3. Lessons learned from building this model can indeed be monetized and many of them may be kept secret.
Yes they can actually.
If you are a shareholder you can either sue (unlikely to succeed) or vote against the board. That's pretty much the only recourse.
Or you could sell or just threaten to sell your shares.
Buying more of the company's shares to take it over is another option.
Good luck doing that with Google.
My current approach is a "model marketplace" (https://accomplice.ai/models) where the most popular open source text-to-image models (VQGAN+CLIP, Disco Diffusion, DALL-E Mega coming soon…), sit alongside the most popular open source style transfer models, and then finally I have the ability for a user to finetune their own models using a simple drag-and-drop tool (https://accomplice.ai/no-code-model-training).
Using this approach a user has enough models to try or train that they can have a higher hit rate. For example, Accomplice currently has finetuned models for photo realistic people (https://accomplice.ai/models/f58bfa91-bb18-406f-a0e1-db00fcf...), watercolor backgrounds (https://accomplice.ai/models/91b8a080-faca-4ff4-8b11-64b0789...), etc…
So theoretically if there were a searchable marketplace of 100s of different finetuned models people could choose from, they would use it much like an iStockPhoto and be able to create the kind of images they want instead of just downloading them.
But it's of course a constant work in progress. Slowly growing though and lots of promising stuff ahead!
What does that link do?
The confirmation link should only be going to accomplice.ai unless Sendgrid is doing some link tracking that I've just forgotten about. Could you forward that email to adam at accomplice dot ai if you get a chance. Thanks for letting me know!
Same thing happened to me multiple times across multiple platforms: SendGrid, Mandrill, MailJet, MailGun. I always turn off the tracking (enabled by default on all of them), but magically its back on a few weeks/months later. I've given up finding a solution and just revisit my settings every few months to check on it.
i.e. The ability to easily take your logo and stylize it: https://accomplice.ai/@adam/iterations/2bcc90ad-3237-486a-8d...
Create a photorealistic avatar whenever you need it: https://accomplice.ai/@adam/iterations/988b7d54-dc39-43b1-b5...
Easily remove the background of a photo: https://accomplice.ai/models/97746c4b-c6f0-49cb-ae1b-859716b...
Upscale a photo: https://accomplice.ai/models/bd4619ee-8202-4cf0-a04e-291820f...
Etc etc. AI can make all this stuff easier. And you have a sense of ownership over what you create. All in one place where you can collaborate on all of it with your team. I feel like that's valuable. It's certainly a tool I've always wanted.
But, also, as a bit of an aside – if the Googles and OpenAIs of the world are just going to bite every artist's style anyway with a mostly black box service and training set… it feels like the option for an artist to train/finetune their own model, promote it and possibly make money off of that is worth trying.
In one sense it's kind of like a much "smarter" photoshop filter, where it can make your own art/photos look more like what you want (ex: Van Gogh, Dali, Picasso, or combinations of those, or something completely weird/new/different).
You could also train the models on your own work and have it generate art in your own style that could inspire you or could be useful to you either as a base to work from or that you could take interesting elements from to create new art.
Similar things can be done in music, by the way, and that would be really useful to musicians too.
Poets could use something like this to create poetry, novel writers to write novels, etc..
This is really an improvement on the collaboration potential between humans and computers -- which is probably why it's called "Accomplice".
The two “african american” ones look south or maybe southeast asian (and the one of those that is a “young...girl” looks like a, maybe young, adult.)
All the ones without a racial/ethnic prompt are white, and disproportionately blue eyed (again, including sclerae.)
(It indicate “diverse”, and yet all of the examples read white or Asian, though the unlabeled darker-skinned male figure in the group of six at the top is ambiguous enough to be plausibly be something else.)
The “beautiful woman with curly red hair” has rather radical facial asymmetry, and straight to slightly wavy hair.
Rule 1 of working in this field, recruit a mid level member of the other company's research group biannually to get the latest gossip.
It's impossible to keep a 1 page or shorter "algorithm" secret, when the creators are geniuses and they hop jobs every year or so.
Fellas with more IQ than games in a baseball season, they just don't forget.
The other reason is that the leaders at Google at the time believed that we would achieve the singularity faster if Jeff Dean periodically sent ideas back 10 years in time to Doug Cutting.
You're saying it's Google who have done this research. In a way that's true. But really it is Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet and Mohammad Norouzi who did it, with material support from Google.
It's likely that some or all of these people would have refused to do the work they do if Google kept it all as their secret sauce.
And moreover, there are excellent reasons why they wouldn't want to. It's not just the obvious that if it all were secret, they wouldn't be able to use it in their non-Google career advancement. It's also that research without the freedom to talk is far more difficult and frustrating.
On paper, scientific papers are supposed to document the whole of the discovery/innovation. So you might think that an insider, who got to read all the secret Google research papers AND all the public ones would have an advantage. But problem is, even the best written papers with full code and comments inevitably leave out things, especially of the "why this and not that" type.
If you're a researcher in the free world, you can just ask. Especially if you have a public track record of great papers yourself, they will WANT to talk to you. You can learn so much more from the interactive process of back and forth questions than you can from a static piece of information like a scientific paper.
If you work for a secretive and command-driven organization, you need to be careful about what you reveal of your own research when you ask. You can't talk freely. The thought of having to justify your communication to some old-school IBM lawyer type, is going to chill even the most enthusiastic reseacher. It's easier to just stay in your own corporate bubble, and focus on the things your corporation does well since at least you can talk freely to your colleagues (although in really paranoid organizations like the NSA or old IBM, even that may not be true). But then at best you specialize, at worst you fall behind.
DALL-E 2 open source implementation - https://news.ycombinator.com/item?id=31228710 - May 2022 (152 comments)
Also:
X-Transformers: A fully-featured transformer with experimental features - https://news.ycombinator.com/item?id=27089208 - May 2021 (37 comments)
Text to Image Generation - https://news.ycombinator.com/item?id=26615791 - March 2021 (88 comments)
Assuming an average training time of 1 week as you said, that gives us about 280k$.
It's also most likely trained for longer than a week, the base model for Dalle-2 was trained for 100-200k GPU hours, so between 2-4x longer than that, we can guess this is roughly similar.
You also never successfully train everything first try, so all in all, to replicate this work just from the paper, we are talking about at least 500k$.
I'll have to read the paper for more details, but it would almost certainly cost less (and take longer) to train a model like this in a more resource constrained situation than Google faces .
The shares would simply be votes towards future training dataset endeavors as no profit would be booked here. Say you buy 5000 out of 500,000 shares, that would give you 1% voting power in what dataset to train.
> 5 on-demand Cloud TPU v3 devices, 5 on-demand Cloud TPU v2 devices, and 100 preemptible Cloud TPU v2 devices for free for 30 days
So up to 7k hours on demand and 70k pre-emptible
Open source is pretty meaningless here
Open source is built on the assumption that you can do more with source code than with binaries. In the case of AI models, the computed weights of models are what's valuable, and the source code used to achieve them is less useful.
> The blog post says 256 GPUs for 2 weeks, so:
> DALL·E would cost $131,604 to train on AWS, assuming a p3.16x-large at market rates. Could be as low as $40k if you already paid for reserved instances.
On compute there is more than enough compute available to open source now via LAION and Eleuther AI to train these models, will just a bit of time.
And medium posts about React hooks and Go generics are gonna be full of “Animated Drake meme but with Rick Astley” kind of fun.
Couldn't you just scrape porn to get copious amount of dataset? Scope would be narrower and thus require less classification and in general it has common themes.
I'm more concerned how expensive it will be to train it on GPU instances. We are looking at A6000s right? That's like $15/hr.
All the complexity wrapped up into one word.
And tags.
Insta has tags.
Plenty of others have nsfw and tags
hint: they are all about one thing and there are a lot of eager volunteers to help on those websites. It would be easy to "normalize/clean/classify" because the pictures would have a consistent theme, thus reducing the amount of parameters.
We are not trying to generate elephants getting railed on a SpaceX rocket flying in oil painting style (although I'm sure theres people into that and its not my place to judge), we are just trying to remove the human cost out of this necessary evil.
I can't believe nobody is investing in "DALL-E-2-4-PORN". This sounds like an X amount of money thrown at a hugely sticky product that can be iterated (with the current trend in hardware) to the point where it literally generates billions of dollars in revenues for ages to come (no pun intended).
Do you prefer known porn actress over non known ones?
I would say that I might know a handful of names but search for them rarly.
They involve real woman (ethical concerns).
They might not show exactly what you are looking for.
Like the good scene is to short or the quality is too bad or there is only one video.
Also variation.
ML porn should be a good thing. Only thing Im not sure is about people creating pedo porn. But even that is better than real pedo porn :|
https://bigscience.notion.site/10743770aae24ff3bdc1b938cf454...
And this is just for a text-only scrape.
For example, in a generative model I'm working on, I have a dataset consisting of ~5M images just blindly scraped from a website. After filtering, this drops down to ~500k images, yet a model trained on that performs worse than one trained on a set of 30k curated images (picked based on a manually evaluated list of "known good" artists), where filtering brings it down to ~18k images. The larger dataset, while containing more information, also contains more errors, many of them pretty hard to filter out.
- lots of people are stimulated by it
- lots of people want DALL-E-2 for porn
- and lots of people are willing to work towards that common goal
The beauty of this is that people are just going to keep coming and coming to it.
Like I'm trying to be mature and serious about this. What's it going to take?
- Community responsible for scraping dataset, generating image dataset from moving pictures, upscaling said dataset.
- Crowdsourced labeling, cleaning, normalizing dataset
- Crowdfunding to train, host, and publish.
All of the above are not easy by any means but its much more achievable than trying to build a generic DALL-E that can create anything.
Really my point is that the scope of the dataset has narrower range in terms of desired output where as DALL-E-2 casts a far wider net.
Specialization is key here.
I guess the primary concern with a porn model would be the ethics of it, which might turn off any company from helping out on training resources (for example, Google's TPU Research Cloud requires you to follow their code of ethics on AI, which would be very difficult to do with a topic as sensitive to people as porn).
There are several potential ways:
The training data may come from unethical sources or processes.
The result may be degrading or demeaning to people who share characteristics with the virtual people depicted (gender, race, etc.)
Increasing the supply may increase the demand, exacerbating the previous issue.
The technology could be trivially repurposed to produce content that is entirely illegal.
That is an oft-overlooked point. In this instance, I think many forget how exploitative pornography actually is. And sex work in general.
We're seeing a lot of garden variety capitalist exploitation happening with ordinary mainstream training data production and labeling work, so I would expect anything adult-adjacent to be correspondingly worse, even though it doesn't have to be.
ಠ_ಠ
Look deepnude is a thing and somebody is making money off it: https://app.deepnude.cc/upload
I don't think the leap is too crazy if we are talking short moving pictures without sound. However, when sound gets involved, this is where it would become very tricky.
I don't think a neural net would have much trouble generating moans in sync to the motion.
Why pay humans for all that fake moaning when an AI could do it?
Given how it renders dog faces, I don't see why it wouldn't be good at human faces too, if trained for it.
One of my more interesting realisations over this development is that being human is apparently a religion to a lot of otherwise secular humans.
If/when general AI comes about, maybe so.
Until then all we've got are a bunch of highly specialized tools/helpers/slaves that may be good (in some sense) at one thing and awful at pretty much everything else.
You could argue that humans themselves are an ensemble of such highly specialized parts that are more than the sum of their parts in some ineffable way that's more of a "we know it when we see it" than something that's formalizable. Machines/computers lack that... so far.
What you describe is already illegal on many jurisdictions.
Are you talking about porn? Or using potential copyright material in an infringing manner? Or perhaps the possibility of "deepfake" porn?
Keep in mind this is a country where if you leave a bad review after you get scammed by someone with evidence, it is defamation. So not quite leadership the world needs in this industry.
Really sad to see ppl on HN flagging all of my comments on this thread. I mean it's not like you can't find celebrity deepfakes including Kpop.
The cat is out of the bag and its only going to get better and faster from here whether some cultures/jurisdictions take offense or not.