Exploring 12M of the 2.3B images used to train Stable Diffusion
waxy.org
waxy.org
Nice tool!
You can also explore the dataset there https://rom1504.github.io/clip-retrieval/
Thanks to approximate knn, it's possible to query and explore that 5B datasets with only 2TB of local storage, anyone can download the knn index and metadata to run that locally too.
Regarding duplicates, indeed it's an interesting topic!
Laion5b deduplicated samples by url+text, but not by image.
To deduplicate by image you need to have an efficient way to compute whether image a and b are the same.
An idea to do that is to compute an hash based on clip embeddings. A further idea would be to train a network actually good at dedup and not only similarity by training on positive and negative pairs, eg with triple loss.
Here's my plan on the topic https://docs.google.com/document/d/1AryWpV0dD_r9x82I_quUzBuR...
If anyone is interested to participate, I'd be happy to guide them to do that. This is an open effort, just join laion discord server and let's talk.
When my prompt isn't working, I often want to check whether the concepts I use are even present in the dataset.
For example, inputting `Jony Ive` returns pictures of Jony Ive in Datasette and pictures of apples and dolls in clip retrieval.
(I know laion 5B is not the same as laion aesthetic 6+, but that's a lesser issue.)
[0] - https://rom1504.github.io/clip-retrieval/
[1] - https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/im...
It works for your example
I guess I'll disable it by default since it seems to confuse people
Using clip for searching is better than direct text indexing for a variety of reasons but here for example because it matches better what stable diffusion sees
Still interesting to have a different view over the dataset!
If you want to scale this out, you could use elastic search
---
Also: There is a joke to be made at Jony's expense regarding the need to turn off aesthetic scoring to see his face.
Quantifying Memorization Across Neural Language Models https://arxiv.org/abs/2202.07646
Deduplicating Training Data Makes Language Models Better https://arxiv.org/abs/2107.06499 https://twitter.com/arankomatsuzaki/status/14154721921003397... https://twitter.com/katherine1ee/status/1415496898241339400
The story is more complex though because the data can often be quite far away from actual neural net training due to preprocessing steps, data augmentations, oversampling settings (it's not uncommon to not sample data uniformly during training), etc. So my favorite place to scrutinize is to build a "batch explorer": During training of the network one dumps batches into pickles immediately before the forward pass of the neural net, then writes a separate explorer that loads the pickles and visualizes them to "see exactly what the neural net sees" during training. Ideally one then spends some quality time (~hours) looking through batches to get a qualitiative sense of what is likely to work or not work and how well. Of course this is also very useful for debugging, as many bugs can be present in the data preprocessing pipeline. But a batch explorer is harder to obtain here because you'd need the full training data/code/settings.
The whole point of this model class is that one can learn one word from one sample, another pixel from another one and so on to master the domain. The emergent, non-trivial generalization is what makes them so fascinating. There is no simple, linear/first order relationship with data and behaviour. Case in point: GPT3 can do few-shot learning despite not having used any explicit few-shot formatted data during training.
Not saying you are wrong, but the story is not as simple as simple supervised learning with small datasets
More details in our repo: https://github.com/simonw/laion-aesthetic-datasette
Search is provided by SQLite FTS5.
sqlite-utils enable-fts data.db images text
https://sqlite-utils.datasette.io/en/stable/cli.html#configu...Writing it out I realize it's not indexed by the attribute I'm retrieving by...
Once it started getting traffic it started running a bit slow, so I bumped it up to a 2 CPU instance with 4GB of RAM and it's been fine since then.
The database file is nearly 4GB and almost all memory is being used, so I guess it all got loaded into RAM by SQLite.
I'll scale it back down again in a few days time, once interest in it wanes a bit.
Presumably it wouldn’t be so hard to hash the images and filter out repeats. Is the idea to keep the duplicates to preserve the description mappings?
It's surprising that these weren't filtered out, and it would be interesting to know the number of unique images. (When it is mentioned that a model was trained on 10 billion images for example, obviously if each image is repeated 5 times then the actual number of images is 2 billions, not 10.)
https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/im...
(Surely there are many more that don't have this specific label).
Oh that's why it is so good at generating Thomas Kinkade style paintings! I ran a bunch of those and they looked pretty good. Some kind of garden cottage prompt with Thomas Kinkade style works very well. Good image consistency with a high success rate, few weird artifacts.
Now I think we don't need the multiverse for that. Give this AI technology a few years and you'll have streaming services a la Netflix where you provide the prompt to create your own movie. What the hell, people will vote "best movie" among the millions submitted by other people. We'll be movie producers like we are nowadays YouTubers. Overabundance of high quality material and so little time to watch them all. Same goes for books, music and everything else that is digital (even software?).
I envision exactly the future as you describe: Feed a song to the AI, it spits out a completely new, whole discography from the artist complete with lyrics and album art that you can listen to infinitely.
"Hey Siri, play me a series about chickens from outer space invading Earth": No problem, here's a 12 hour marathon, complete with a coherent storyline, plot twists, good acting and voice lines.
The only thing that is currently limiting us is computing power, and given enough time, the barrier will be overcome.
A human brain is just a series of inputs, a function that transforms them, and a series of outputs.
Don't know if this is sarcasm, if it is, ignore the rest of the comment.
Honestly, it sounds terrible.
Good shows are well written, coherent and, most of all, narrow in scope.
If an AI can write the next Better Call Saul, great.
Randomly patching together tropes sounds more like kids drawings, that are interesting on an artistic point of view, maybe, given their limited knowledge of reality and narrative, but terribly boring and confusing as a form of entertainment.
Unless the audience is kids, they love that stuff, for reasons we don't understand anymore as we grow up.
Endless entertainment it's already there, you can't watch it simply because most of the content you're talking about is not on air and streaming platforms don't buy it, because it's shit.
Not that I don't like shit, I've watched more troma/SyFi/random low budget Asian movies (I am a huge fan of ninja movies) than necessary, but if we stop to think that there are already 6 or 7 sharknado movies (that are exactly the kind of endless entertainment you talk about, they are probably generated in some way), maybe it's not the volume of content that's missing, but probably content that's worth watching.
Example?
I don't think I've ever wanted to watch something that only a computer could generate.
I don't have any trouble finding youtube channels that I like to watch and ignoring the rest, and I suspect I won't have any trouble finding movies generated using AI as a production tool that I want to watch either.
as per HN rule, don't ask this question.
> Parent post was talking about assisted movie generation where a human is making a movie and using the AI as a tool to make the content
which is exactly the problem we don't have.
there are thousands of scripts written everyday that never see the green light.
> and it will lead to a creative revolution
it won't
main reason why content is not produced is money.
unless you find a way to create an infinite supply of money and an infinite amount of paying audience for that content, more content is a problem, not a solution.
> I don't have any trouble finding youtube channels that I like to watch and ignoring the rest,
so what's the problem?
There's already infinite content out there, what does the "AI" brings to the table that will make any difference, other than marketing, like 3D movies?
Have you watched any 3D movie recently?
"there are thousands of scripts written everyday that never see the green light." "main reason why content is not produced is money."
So if there are plenty of ideas and not enough money, and you could put those ideas into a box and spit out a movie that would normally cost millions, that's good right?
big chunk of the budget is spent on marketing.
if you produce something that nobody watch, it's like the sound of the tree falling where nobody can hear it.
if you know how to use AI to cut that cost, I'm all ears.
Also: Al Pacino will want his money if you use his name, even if he is not actually acting in the movie.
Reality is that there's plenty of ideas, true, that would not make any money though.
Studios don't like to work at loss.
Rick and Morty costs 1.5 million dollars per episode and from what we've heard from director Erica Hayes, a single episode takes anywhere between 9 and 12 months to create from ideation to completion
Cutting costs seems to be the main reason AI is being explored. If you go to studio asking for budget to create a movie and predict "10 000 people will watch it", they will laugh in your face. If one person with the help of AI can make the movie and 10 000 people will watch it, its a win for everyone involved.
I dont see youtube channels having enormous budgets for marketing, yet they find sizeable audience and make profit still. Once you lower the cost of production, you dont need huge marketing budgets to secure profits.
Because they mainly support one person.
You don't need a big budget to sell lemonade on the street, you can make a salary out of it, doesn't mean you have become a tycoon or have revolutionized the lemonade stand industry.
> I dont see youtube channels having enormous budgets for marketing
Have you seen those ADS every 15 seconds?
That's the marketing budget, the whole YouTube ADS revenue is the marketing budget.
A social media account with a couple million followers i.e. https://www.instagram.com/lilmiquela/?hl=en
> Also: Al Pacino will want his money if you use his name, even if he is not actually acting in the movie.
The creator doesn't need Al Pacino. I'm not following this one. Rick and Morty doesn't need Angelina Jolie
> Reality is that there's plenty of ideas, true, that would not make any money though.
Plenty of websites, 99.999% of them don't make any money. I definitely still think the easier it is to make a website the better.
> Rick and Morty costs 1.5 million dollars per episode and from what we've heard from director Erica Hayes, a single episode takes anywhere between 9 and 12 months to create from ideation to completion
If that includes marketing and advertising, that's amazingly cheap. Cut the creation time down to a few weeks, throw in a few product placements, post it on your social media/youtube/etc and you have your own movie studio.
Not really the same - there is a range from good to bad on youtube, because real people are adding the creative spark. There is no reason to suspect AI will generate such a range, and its unclear we will ever get to the point where AI can do "creativity" by itself.
> randomly patched together tropes
Sounds like a dream, not as in what I wish for, but what I experience at night.
So, if you think about it like a tool to enable a form of lucid dreaming, it may be something interesting.
Of course you have to find a way to get for your brain what you want to see in "real time", but I think we will get there.
we usually call that tool psychedelic drugs.
there are devices being developed for that purpose, I don't think they will ever be reliable, AI is not necessary for that.
On the philosophical implications of lucid dreams
http://geekdommovies.com/heres-what-the-live-action-lion-kin...
I have a decently old LG 3d tv that can actually turn 2d into 3d, and its actually a lot of fun to watch certain stuff in 3d mode.
txt2img is quite limited and img2img is really where the power is. With a little intermittent guidance hand of a human.
What took 100s or 1000s of people to write, act, record, post process, Better Call Saul might be doable by a a team 1/100th the size possibly even to a single individual. Which means while it might not just instantly spit it out. I'd just like youtube, there will be an incredible amount of great content to watch, far more than anyone could ever realistically watch.
And of course there will be lots of utter trash as well.
But if it took 100 people to make better call saul and now 100 individuals can make 100x different "better call sauces"
Quick reminder, there are infinitely many even numbers and none of them are odd.
A given infinite (or transfinite) set does not necessarily contain all imaginable elements.
Is there a universe with channels focusing on these subjects:
"Video of the last tears of terminal cancer patients as Jerry Lewis tells jokes about his dick"
"This guy doesn't like ice cream but he eats it to reassure his girlfriend that he isn't vegan"
"A single hair grows on a scalp"
"Infant children read two-century-old stories about shopping for goose-grease and comment on the prosody"
"Gameshows where an entire country's left-handed population guesses how you'll die"
This is the whole point of the Rick and Morty cable bit. There are things that would not be on TV in any universe that invents TV. It's hilarious to pretend they would be.
In a show, there are a certain number of frames. In one frame, there are a certain number of pixels. Each pixel can be one of some number of colors. An infinite TV would be able to show every combination of every color of pixels, follow by every combination of frames, simultaneously and forever. All those shows are in there. Not only that, but all of this is also countably infinite.
In a multiverse situation you're watching actual content from an infinity of universes where beings exist who have produced and selected that content.
You're not watching an infinite amount of static and magically selecting the parts of static that are coincidentally equivalent to specific content.
All multiverse content must be watchable and creatable by the kind of creature that creates a television. So content that is unwatchable or uncreatable by any conceivable creature will not be in that infinite set.
It is very easy to describe impossible content and any impossible content will not be on the multiverse TV.
Trivial counterexamples that are describable but uncreatable, in any universe similar enough to ours that it has television:
-a channel that reruns actual filmed footage of a given universe's big-bang will not exist.
-a channel that shows *only* accurate, continuous footage of the future of that universe.
-a channel that shows the result of dividing by zero.
Channels that may or may not be uncreatable:
-a channel that induces in the viewer the sensation of smelling their mother's hair.
-a channel that causes any viewer to eat their own feet.
-a channel that cures human retinal cancer.
-a channel that shows you what you are doing, right now, in your own home in this universe, like a mirror. Note that this requires some connection between our universe and the other universe and there's no guarantee, in a multiverse situation, that connections between the universes are also a complete graph.
These examples are more important. We know they are either possible or impossible but we do not know which. Just saying "multiverses are infinite" doesn't answer the question.
For further reading, review https://en.wikipedia.org/wiki/Absolute_Infinite and remember that a channel is a defined set.
that's the movie GAN, more interesting would be the zero-shot translations of epic foreign/historic films into your native culture
There are quite a few books written, so maybe transfer learning from that?
None, according to the MPAA.
Imagine uncovering a movie that was never made but featured actors you know. If Steven Spielberg can make the movie, then there is an undiscovered sequence of 1's and 0's that already is that movie, a sequence that could be discovered without actually making the movie. Imagine "mining for movies"...
Of course that sequence is likely impossible to ever predict enough of to actually discover something real... but it's a fun thought experiment.
https://screenrant.com/rick-morty-interdimensional-cable-epi...
It's extremely hard to make good content. Teams of extremely skilled, well-paid people, even ones who have succeeded before, fail regularly. And that's with complicated filtering mechanisms and review cycles to limit who has access and keep the worst of it from getting out.
But not everybody needs everything to be actually good. My partner will sometimes unwind on a Friday by watching bad action movies; the bad ones are in some ways better, as they require less work to understand and are amusing in their own way. Or there's a mobile game I play when I want to not think, where you have to conquer a graph of nodes. The levels are clearly auto-generated, and it's fine.
I think that kind of serviceable junk is where we might see AI get to in a couple of decades, made for an audience that will get a weed gummy and a six pack and ask for a "sci-fi action adventure with lots of explosions" and get something with a half-assed plot, forgettable stereotypical characters and visuals that don't totally make sense, but that's fine. You won't learn anything, you won't be particularly moved, and you won't ever watch it again, but it will be a perfectly cromulent distraction between clocking out and going to bed.
We are going to absolutely drown in crap. Just as we do now, it's just going to flood the internet at an unimaginable pace. We'll probably train AIs to help us find stuff and tell what's real/true/etc; it's going to be a heck of an arms race.
It's going to be one hell of an arms race.
Plus you need to encode the ideas of plots/characters/scenes, and have through-lines that go multiple hours. It seems like with the current kit it’s hard to even make a consistent looking children’s book with hand picked illustrations.
My gut is we are more than a few years off, but maybe I’m underestimating the low hanging fruit?
If a tree falls down in a forest, and there is no one there to hear it. Does it make a sound?
If you can generate infinite material, how do you judge quality?
You're extrapolating an idea based on what movies are, fundamentally. But you don't take in consideration what movies are not. Watching a movie is also a social experience. Going to the movie theater, waiting years for a big blockbuster title, watching something with friends. Word of mouth recommendation is a very big thing. If a close friend recommends me something (be it a movie or a book), I'm much more inclined to like it just for the human connection it provides (reading or watching something other people enjoyed is a means of accessing someone else's psyche).
If every time you watch a movie you have the knowledge that there is a movie that is slightly better a prompt away, why bother finishing this one? If you know you probably won't finish the movie you generated, why bother starting one? So what do you do? You end up rewatching The Office.
Sure, if you tell me this will be possible in a couple of years, I won't object. The point is: will you pay for it on a recurring basis? Because if you don't, this will be no more than a very cool tech project.
----
I've recently had this idea for a sci-fi book: in a future not so distant, society is divided between tech and non-tech people. Tech people created pretty much everything they said they would create. AGI, smarter-than-human robots, you name it. But, it didn't change society at all. Companies still employ regular humans, people still watch regular made-by-human movies and eat handmade pizzas and drink their human-made lattes in hipster coffeshops. So tech people are naturally very frustrated at non-tech people, because they're not using optmizing their lives and business enough. And then you have this awkward situation where you have all these robots with brains of size of a galaxy lying around, doing nothing. And then some of them start developing depression, for spending too much time idle. And then the tech people have to rush to develop psychiatric robots. And then some robots decide to unionize, and others start writing books about how humans are taking jobs that were supposed to be automated.
Self awareness here would have lead to the removal of "enormously popular".
I'm starting to realize Stable Diffusion doesn't understand many words, but it's hard to tell which words are causing it problems when engineering a prompt. Searching this dataset for a term is a great way to tell whether Stable Diffusion is likely to "understand" what I mean when I say that term; if there are few results, or if the results aren't really representative of what I mean, Stable Diffusion is likely to produce garbage outputs for those terms.
http://laion-aesthetic.datasette.io/laion-aesthetic-6pls/ima...
> There is a certain degree of duplication because we used URL+text as deduplication criteria. The same image with the same caption may sit at different URLs, causing duplicates. The same image with other captions is not, however, considered duplicated.
I am surprised that image-to-image dupes aren't removed, though, as the cosine similarity trick the page mentions would work for that too.
How could one go about deduping images? Maybe using something similar to rsync protocol? Cheap hash method, then a more expensive one, then a full comparison, maybe. Even so 2B+ images... and you are talking about saving on storage costs, mostly which is quite cheap these days.
I think the way you do it is to train a model to represent images as vectors. Then you put those vectors into a BTree which will allow you to efficiently query for the "nearest neighbor" to an image on log(n) time. You calibrate to find a distance that picks up duplicates without getting too many non-duplicates and then it's n log(n) time rather than n^2.
If that's still too slow there is also a thing called ANNOY which lets you do approximate nearest neighbor faster.
It's performant enough even at scale.
If you need to do something closer to pairwise (for instance, because you can't make a cheap hash of images which papers over differences in compression), make the hash table for the text descriptions, then compare the images within buckets. Of the few 5 or 6 text fields I just spot checked (not even close to random selection) the worst false positive I found (in the 12M data set) was 3 pairs of two duplicates with the same description. On the other hand I found one set of 76 identical images with the same description.
Of course actual hash algorithms are a bit cleverer, there are a number to choose from depending on what you want to consider a duplicate (cropping, flips, rotations, etc)
That approach works well when the images are basically the same. It doesn't work so well when you're trying to find images which are either different photos of the same subject or where one of them is a crop of a larger image or has been modified more heavily. A number of years back I used OpenCV for that task[1] to identify the source of a given thumbnail image in a larger master file and used phash to validate that a new higher resolution thumbnail was highly similar to the original low-res thumbnail after trying to match the original crop & rotation. I imagine there are far more sophisticated tools for that now but at the time phash felt basically free in comparison the amount of computation which OpenCV required.
1. https://blogs.loc.gov/thesignal/2014/08/upgrading-image-thum...
You can try this on images.yandex.com - they do similarity search with embeddings. Upload any photo and you'll get millions of similar photos, unlike Google that has only exact duplicate search. It's diverse like Pinterest but without the logins.
Query image: https://cdn.discordapp.com/attachments/1005626182869467157/1...
Yandex similarity search results: https://yandex.com/images/search?rpt=imageview&url=https%3A%...
this is too inflammatory in my opinion.
- some works are in public domain
- tech companies have profited from creators for a long time. I'm sure some arrangement could be made for profit sharing for artists who care about money, but it's too early for that (no profits, I'm sure most companies are losing money on AI art)
- some artists care about art or fame more than money. Their art will not be devalued by AI, if anything, constant usage of their names in prompts is going to make them massively popular and direct people to source material or merch, which they may buy.
- some artists are dead and don't care anymore. Their "estate" vulnerable to takeovers by "long lost but recently found" relatives, who don't care about art itself, only about money. Many such stories.
One example, albeit not in paintings but in music is Jimi Hendrix Estate. They used to do copyright strikes on YouTube in order to remove fan-made compilations of rare material (cleaned up sound of live concerts, multiple sources mixed into one etc.), without intentions to ever release an alternative.
Would one need to retrain the entire dataset? Or is there typically a way to just add an incremental batch?
For example Stable Diffusion knows much better than MidJourney what a cat looks like, MidJourney knows what a Hacker Cat looks like, while Stable Diffusion doesn't (you can tell it to make a cat in a hoodie in front of a laptop, but it won't come up with that on its own). Meanwhile for landscapes Stable Diffusion seems to have no problem with imagination. How much of that is simply due to blindspots in the training data?
Also, a UI issue: the sorting arrow feels wrong:
https://i.imgur.com/uyUXAXy.png
The norm is that when the arrow is pointing down, the data are currently sorting descendingly (I'm aware you can interpret it as "what will happen if you click it", but the norm is to show what currently is using.)
When I designed that feature I looked at a bunch of systems and found examples of arrows in both directions.
> “realistic 3d rendering of mickey mouse working on a vintage computer doing his taxes” on DALL·E 2 (left) vs. Stable Diffusion (right)
Well, but the mickey mouse in the right isn't "realistic", or even 3D. It's straight up just a 2D Mickey image pasted there.
Sure, the achievements of ML models lately are impressive, but it's so slow in learning. We are brute forcing the DNNs it feels to me, which is not something that smacks of great achievement.
You and I have never seen even 100,000 photos in our lives. Well, maybe the video stream from our eyes is a little different. But it's not a billion fundamentally different images.
Is there anything I can read about why it is so slow to learn? How will it ever get faster? What next jump will fix this, or what am I missing as a lay person?
That's where you start with an existing large model, and train a new model on top of it by feeding in new images.
What's fascinating about transfer learning is that you don't need to give it a lot of new images, at all. Just a few hundred extras can create a model that's frighteningly accurate for tasks like image labeling.
This is pretty much how all AI models work today. Take a look at the Stable Diffusion model card: https://github.com/CompVis/stable-diffusion/blob/main/Stable...
They ran multiple training sessions with progressively smaller (and higher quality) images to get the final result.
* We have seen more than 100,000 "photos" in the sense that photos are just images - if photos are just images, we have a constant feed of "photos" every single moment our eyes are open. Of course, that's not the same as these training datasets, but it is still worth keeping in mind.
* All of these things trained on massive datasets with self-supervised learning are in a sense addressing the "slowness" of learning you mention, since self-supervised (aka no annotations are needed beyond the data itself) "pre-training" on the massive datasets can then enable training for downstream tasks with way less data.
* Arguably requiring massive datasets for pre-training is still a bit lame, but then again the 4-5 years of life it takes to reach pretty advanced intelligence in humans represents a whoooole lot of data. And as with self-supervised learning on these massive models, a lot of intelligence seems to come down to learning to predict the future from sensory input.
* Humans also come with a lot of pre-wiring done by evolution, whereas these models are trained from scratch. Evolutionary wiring represents its own sort of "pre-training", of course.
So basically, it is not so slow to learn as it seems. Arguably it could get faster once we train multimodal models and concepts from text can reinforce learning to understand images and so on, and people are working on it (eg GATO). There may also need to be a separation between low level 'instinct' intelligence and high-level 'reasoning' intelligence; AI still sucks at the second one.
https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/im...
Clearly quality of labeling isn't nearly as important once you are training on billions of images.
A cool fact is that this model fit ~5B images in 900M parameter model which is tiny compared to the size of the data.
People need to get over the metaphors. If you want to spend your time learn about the mathematics under the hood, there will be less "mysteries" then.
If someone's 30, that'd only require seeing 10 images a day. For most people that quota is probably fulfilled within a couple of minutes of watching TV or browsing social media, even if the video stream from our eyes otherwise counts as nothing.
We've also had about 4 billion years of evolution, slowly adjusting our genome with an unfathomable amount of data. Gradient descent is blazing fast by comparison.
I would argue precisely the opposite (as you alude to), it's more than 100s of billions of fundamentally (what does this even mean?) different images. Calculate the frequency at which your eyes sample, think of the times the angle changes (new images), multiple by your age, multiple by 2 for two eyes looking slightly different directions, factor in your the noise your brain has in forming the image "in your head" because you drank too much... you can continue adding factors (hours of "TV" watched on average) add-naseum.
It seems that "slow to learn" has a "real" target/bound, what humans are capable of. If it takes Bob Ross decades to paint all the "images" in his head, then maybe we should go easy on the algorithims?
Active learning is a good way to improve sample efficiency - however, as others note, don't underestimate the quantity of learning data that a human baby needs for certain skills even with good, evolution-optimized priors.
Are they ok with their stock photos being used to train a service that's likely to bite into their stock photo business?
An artist decides to sell prints of an expensive artwork, he or she publishes a photo on their website. AI scrapers get the images in data set update. Game over for the artist.
I hope for a class action over this training data sets. I get that kids have fun with the new photoshop filters. I get that software is eating "the world" but someone must wake up an push the kill switch. It is possible.
At around the same time that it actually resembles "the new data", or maybe even shares a single quality with it? Being a physical object, it is bound by physical properties... like scarcity. Data, being an abstract concept, suffers no such constraint. Same story for whatever artistic technique you've imagined to be not only valuable, but novel to all of human experience and wholly owned by you and you alone. Your valuation of your worth and that of your labor is laughably overinflated and the market is telling you so. Period.
Good luck with that. :)
Everything around you is a product of Intellectual Property. Your logic is laughably naive.
The market will shift to AI generated nonsense moved by greed and over-optimization. People will reject this en masse.
And there is one big undeniable reason: Human Psychology.
This makes me vaguely uneasy. All these models and tools are almost exclusively "western".
Plenty of asian / artists / styles on the datasets no?
Personally, I’m not even getting out my popcorn yet.
…and they can still lose it all to typos or fees permanently.
This is clearly different. The value has been demonstrated - and it has clear implications for a lot of jobs.
Just just just starting to scratch the surface. 2% of the data gathered, sources identified. These are a couple of sites we can now source as the primary powerers of AI. Barely reviewed, dived into in terms of the content itself. We have so little sense & appreciation for what lurks beneath, but this is a go.
In a world where the lay public didn’t really know about photoshop, photoshop would be a terrifying weapon.
Likewise modern ML is for the most part mysterious and/or menacing because it’s opaque and arcane and mostly controlled by big corporate R&D labs.
Get some charismatic science popularizers out there teaching people how it works, all the sudden not such a big scary thing.