Generating Fashion Using AI
twitter.com
twitter.com
It feels like everyone's so desperate to create or predict the next AI unicorn, that no one's paying attention to the fundamentals. Gives me weird dotcom bubble vibes.
I still cannot see how this can fundamentally change the fashion industry. Fashion is not about design, it's about generating desire to buy. You hire models and influencers to showcase your items, so people think they can be just like them, you just have buy the same piece of clothing. AI doesn't change that. You could totally have a niche for AI-generated clothing, but that's it.
Besides, clothes are a physical item. AI can automate the generation of a "blueprint", but it is still heavily constrained by the physical result. It would be mostly hit and miss. Feels more efficient to have designer use their experience to create sketches that would actually look good in real life.
All in all, I am amazed as much as I am skeptical of the entire synthetic art revolution going on. It's been what? A month? It's too soon. Aren't we jumping the gun?
Whatever happened to boring opinions? "Hmm that's amazing, but we'll see what happens, it's too soon". I'm yet to be convinced of how, beyond being an overall better tool, synthetic art will fundamentally change businesses.
Nobody presented it as changing the fashion industry. It's generating new outfits in a video, and that's super cool.
I don't think anybody's moving too fast or missing fundamentals here, at least not yet. This is all brand-new, so people are just having a fun time exploring.
The most obvious application of this is to replace, augment or make more economic live action computer graphics techniques.
That could actually be pretty interesting though, liberating even, because it would open up a previously incredibly expensive field... giving more access to compete with CG previously only available to high budget hollywood. Imagine if manipulating a live action scene was as easy as dropping some descriptions into some program and paying for a little compute.
The influencers can also be generated by AI (or even just 3D modelling), since so much of the marketing is digital. And it's already happening (mostly in asian countries):
https://www.dexerto.com/entertainment/ai-created-influencer-...
These videos are just demonstarting the concept design part created by something not even designed for this specific use case. It's true that making a product from a concept design is another problem. But it's not hard to imagine that AI could be trained specifically for this on actual design plans and using physically accurate cloth simulation in the training loop to generate something that could be actually built. After all if AI can improve protein folding predictions why not cloth folding :). (OK, it' not completely the same but both are about 3D structure).
We already have some pretty good sub-millimeter level cloth simulations (not using AI but could help AI): https://www.youtube.com/watch?v=Mrdkyv0yXxY
And of course people are excited when they have a new tool they can use, they try to find more usecases for it and some will work out some not. In this case they even have the open-source blueprint for the tool so they can finetune it for their own ideas.
One big issue is that big companies already with some regularity rip off boutique designs and then undercut them. Big data and harnessing AI will mostly make that situation worse.
Examples from Shein, a Chinese fast fashion retailer that already uses strategies like this: https://www.dazeddigital.com/fashion/article/55146/1/shein-f...
What exactly are the fundamentals?
The transformer revolution is pretty foundational and the inability to see the societal impact of this technology can only be explained by a lack of imagination.
Majority of peoples time is spent in digital worlds, the ability to create synthetic images, text, etc. is going to radically alter the way we spend our time online. Both malicious and marvelous use cases will be discovered, but to hand-wave it all away as a "Dotcom bubble vibe" is tech cynicism at its worst.
That's just the bias for people in your bubble, the majority of the population of this planet spends its time overwhelmingly in the real world. Calling it "digital worlds" is also a big indication of tech bro bias.
AI does point us in an exciting new direction though. This stuff is baby steps and currently at day 1 of germinating a whole new kind of fruit 20yrs down the line. And that's all it needs to do, to be revolutionary. We need new tools and directions.
True automation is coming to art.
Mindblowing.
https://twitter.com/remi_molettee/status/1564632028959629319
https://twitter.com/matthen2/status/1564192185909739521
https://twitter.com/Infinite__Vibes/status/15650454342293340...
https://twitter.com/thibaudz/status/1564892979789045760
https://twitter.com/chav_ez/status/1565806042344087552
https://twitter.com/zippy731/status/1564616100477820938
https://twitter.com/RonnyKhalil/status/1565024524181159941
https://twitter.com/remi_molettee/status/1565356181190807553
https://twitter.com/remi_molettee/status/1563187170734927872
https://twitter.com/replicatehq/status/1564354673108127744
https://twitter.com/genekogan/status/1564626995979370505
https://twitter.com/ala_art_lab/status/1565984951178346496
https://twitter.com/DrewMedina20/status/1565746320953966592
https://twitter.com/Aiartitune/status/1565795049102786560
https://twitter.com/Aiartitune/status/1563651144806645760
https://twitter.com/ala_art_lab/status/1565603328003870720
https://twitter.com/pharmapsychotic/status/15642809223625973...
https://twitter.com/Aiartitune/status/1563832168119517186
https://twitter.com/socalpathy/status/1565899540451966977
https://twitter.com/zippy731/status/1565342075196870656
https://twitter.com/makeitrad1/status/1563335226524282882
https://twitter.com/mrflosunday/status/1565885053753761792
https://twitter.com/TomLikesRobots/status/156488734249359769...
https://twitter.com/EuclideanPlane/status/156421740831482675...
https://twitter.com/erocdrahs/status/1565320455162044417
https://twitter.com/Carl_Ingram_art/status/15630745562893967...
https://twitter.com/benscottpye/status/1565352548608778242
https://twitter.com/MichaelCarychao/status/15645904797940613...
https://twitter.com/AiJoe_eth/status/1564221320916779011
https://twitter.com/Carl_Ingram_art/status/15646973666696273...
https://twitter.com/ChekhovEugene/status/1565880769477738497
https://twitter.com/originalmaderix/status/15656282243520552...
https://twitter.com/Aiartitune/status/1564177213888643072
https://twitter.com/Infinite__Vibes/status/15646387276616581...
Has someone a good intro to stable diffusion for someone who doesn't know a lot about AI?
Couldn’t find any links on YT. Please share if anyone has them.
There’s also the problem that AI can’t be specific. I can’t design merch with a specific video game logo or band name. The output always has that “AI dream residue”.
This is a useful tool only for creative inspiration.
At least with the photorealistic prompts, there is always this feeling of "this doesn't look right". Right now it's easy to pinpoint what it is. Usually is a completed distorted face. But they might get better at this. But probably there will always be that "uneasy feeling".
I would like to understand if there's a more automated way of doing this.
I don't think that's true. Just like Gimp can be used to make better results than Photoshop (or even mspaint), the quality of the results is up to the artist/user, not the tool itself.
Some SD outputs are truly amazing, but so is some DALL-E 2 outputs. Maybe DALL-E 2 is easier to use for beginners, but people who are good with prompts will get as good output with SD as with DALL-E.
I agree it might do a the job decently for a dress in this instance.
In addition to linguistic precision, the parameters involved in prompt composition, for a perfectly controlled artistic result, require technical knowledge, a sense of style and historical knowledge. The more related keywords involved in the composition, the greater the artist's control over the final result. Example: the prompt
_A distant futuristic city full of tall buildings inside a huge transparent glass dome, In the middle of an arid desert full of big dunes, Sunbeams, Artstation, Dark sky full of stars with a bright sun, Massive scale, Fog, Very Detailed, Cinematic, Colorful_
is more sophisticated than just
_A city full of tall buildings inside a huge transparent glass dome_
Note that the conceptual density, hence the quality, of the prompt depends heavily on the cultural and linguistic background of the person composing. In fact, a quality prompt is very similar to a movie scene described in a script/storyboard [by the way, there go the Production Designers, along with the concept artists, graphic designers, set designers, costume designers, lighting designers… ].
In an attempt to monetize the fruits of the new technology, Internet entrepreneurs will be forced by the invisible hand of the job market to delve deeper into language skills. It will be a benign side effect, I think, considering the current state of the Internet. Perhaps this will lead to a better articulation of ideas in the network environment.
Just as YouTube influencers have a knack for dealing with the visual aspects of human interactions, aspiring prompt engineers will have to excel at sniffing out the nuances of human expression. They have great potential to be the new cool professionals of the digital economy, as were web designers, and later influencers — who, with the end of social networks, now tend to lose relevance.
To differentiate themselves, prompt engineers will have to be avid readers and practitioners of semiotics/semiology.
Umberto Eco and the structuralists may return to fashion.
(*) I used a prompt by Simon Willison
So they have to be a character out of a William Gibson novel? Do they rip the labels off their clothes too?
PS: This is awesome.
You are assuming that the models themselves respond accurately to linguistic clues. Actually they embody the cloud of random noise, prejudices, innacuracies and misconceptions in the training data and then pile on a big layer of extra noise by virtue of their stochastic nature.
So this isn't a case of the learned academic with extensive domain knowledge steering a precision machine. It's more like a someone poking a huge chaotic furnace with a stick and seeing what comes out.
Imagine someone saying that about Photoshop when it was first introduced...
"Why would you even use Photoshop if so you still have learn a tool in order to use it? Might as well just use a canvas with colors... The whole point of Photoshop is that your grandma should be able to use it!"
No, every tool in the world is not meant to make it zero effort to do something. Some tools are meant to incrementally make it easier, or even just make it less effort for people who already are good at something. And this is OK.
There's a huge difference between Gimp/Photoshop and an image generation model.
If a particular model can't generate faces properly then the "artist/user" can't get around that unless they develop a new model or find one that can fix the output of the first.
Have to confess I haven't been bothered to log into D2 since the SD beta started. Which is crazy because my mind was blown when it first launched back in April, but it for now seems like closed AI simply can't keep up with the open source community swarm.
OpenAI is just another example of a typical SaaS business misusing the word to make themselves seem "nicer" than they are. Ultimately, it's a for-profit business and it will operate as such, shouldn't surprise anybody.
If you want a SaaS model, you should have all of the things that the SAAS model supports including premium pricing and enterprise tiers, OpenAI needs to get their act together on DallE2 as no serious business use case will use the consumer pricing and hacked together unofficial apis.
For me, if you offer something software-y behind payment, and the software can only be accessed online, it's a SaaS.
No need to offer premium pricing, enterprise tiers, support or anything else. Putting up a online API/UI that is locked behind payment, it's a SaaS (in my eyes).
I suspect there will be a lot of companies that build businesses out of chaining models together for specific kinds of users. The more you can focus on a niche, the more you can paper over some of the limitations in the models.
There is Twitter too https://twitter.com/UnshushProject
I started with a lot of enthusiasm but discovered you need to be A LOT more social (or lucky) to make people see things!
Next time maybe with some code and a Show HN it will be more fun.
Standing infront of your augmented reality mirror in the evening swiping through AI generated outfits for the next day that are made automatically during the night and shipped to you by drone the next morning to wear for the day and then dispose of.
An environmental catastrophe with today's technology.
But if we can fix that, say make the clothes with a 3d printer out of yesterday's recycled clothes, with energy from renewable sources, then that sounds pretty cool actually.
Adding a bit more imagination to your initial idea make the entire idea actually a net-positive. People can start recycle clothes (good for the planet), with a new design everyday (good for people who like that) and it provides a business (good for the economy).
You wear AR glasses and can choose whatever virtual fashion you want. Other people wearing glasses too can also see your outfit. So you make something physical into something virtual, thereby not having fast fashion at all.
So, my non-artificial intelligence has generated the following new fashion: Last year's.
(*) - Note I'm not saying change clothes on themselves, but the set of clothes they own/use.