Break-a-Scene: Extracting Multiple Concepts from a Single Image
omriavrahami.com
omriavrahami.com
Make you can become an independent filmmaker. It seems like creatives won't need Hollywood tools or capital in the future, either.
Musings aside, I'm not so certain people won't need a photographer to be present at weddings anyway. They can't exactly trust that to an appliance. They get one shot to take the pictures, and they need to make it work.
So if you had taken a few of the low quality smartphone pictures that guests took (had smart phones been as ubiquitous when I was married) and a beautiful background of our wedding location, and asked the AI to create some gorgeous portraits of us, with just the right "golden hour" light and the right look of love in our eyes, would the final portraits tucked away in our album be any different? Would they even be better?
It reminds me of the crowds of people in front of the Mona Lisa with their cameras out, so they can own their own, personal, blurry picture of the world's most photographed piece of art, instead of buying the perfectly-lit postcard in the gift shop. What's actually different about that much-ridiculed desire, vs. wanting to have the "authentic" photographs that that the photographer actually took (and then photoshopped, etc.).
That's why I don't think people would like IA to improve their photos of their weddings, or so on. People don't make a wedding to have the perfect photo at sunset, but to have a nice event with people they love, and the photos are to capture that and nothing else. People's favourite photos (and not only from weddigns) are blurry, caught with friends, etc., because those pictures capture the feeling of a good moment.
There's a difference between discovering an old photo and thinking "haha I don't remember taking this one, we looked so good" and seeing it and thinking "haha I don't remember taking this, oh right because it didn't happen". In the second case, there's no reason to even have a photo album or generate anything at wedding-time. Just store the single base photo, and then 50 years later when you want to show the grandkids you can just have the computer generate whatever they want to see.
When I have seen an excellent wedding photographer, the skill seems to be 20% photography and 80% people management. Great wedding photographers seem to be some sort of strange combination of entertainer, project-manager and therapist (who also takes photos).
Construct new scenes anytime and algorithmically.
"Generate pictures of me with each of my friends, making crazy faces, while drinking mai tai's at my favorite tropical resort bar."
Make real life, "I am so rich and happy with my material wealth and vast travel budget" influencers redundant.
However, I am not looking forward to Facebook spamming friends with ads posing as personal recommendations, showing me at my house happily using sketchy products or butt bleacher.
They will be sneaky: Show me, but with my face obscured in some way. But recognizable. But deniable. But recognizable.
Actually, if the software is really smart (and maybe supplied by Google) you could ask "remove all of my kids' spouses that will be divorced at some point in the future"
or until the VC money dries out and when it does you will have two weeks to download it all.
We won't be able to value photography in the same way, so the place photography has in society will naturally change as a result.
If all photography because meaning-neutral, does that mean we'll value the medium in the same way we value stock photography?
Time to collect ideas for that avant-garde fine art photoshoot I can now do without a camera, props, site searches, or expensive supermodels.
I have also been collecting and developing story ideas. We are not that far away from one person being able to create an entire movie or TV series.
This may lead to some really great movies. Having every detail coming from and integrated in one mind, with one vision, iterating without friction, will be unprecedented. No need for actors and their schedules, retakes, wardrobe, sets, caterers or VIP RV's.
Or if models have long term memory at that stage, two "minds"!
Obviously, video is going to take more processing power. But it is inevitable.
Two components I need: (1) scenes that stay consistent, and true to natural lighting and physics, and (2) the ability to create characters whose proportions, look, facial expressions, voice and style of movement are consistent by default, but easily adjusted.
We are entering the singularity. This isn't stopping.
Actors don't just mindlessly repeat the lines some singular genius wrote. A lot of famous movie lines/scenes were improvised by the actors.
The benefits of isolation do not negate the benefits of collaboration.
The difference will be that both ends of the spectrum will be possible.
Individuals will be able to do it all.
But people will continue to be able to collaborate with anyone. Co-writers, actors, set designers, etc.
An actor who brings a lot could now inhabit several characters in the same story. Or every character!
A creative set designer is no longer limited to found or built sets.
Starting with one writer, there will be no limits on tasks they can do, or who they collaborate with, and how they divide up the work.
Collaborations of all kinds will continue.
But teams will generally be much tighter.
Add a chatGPT like interface where you can communicate by sentences, sprinkle in a bit of speech recognition. And you have computers from 90s sci fi movies. "Computer, give this cat an hawaiian shirt, and make it surf."
The future is now.
Life imitating art imitating life. Exciting times indeed.
this singularity is composed of many elements (themes, forms, subjects etc) and sub-elements. What I would have loved to see in this paper is a means by which the heirachies of these elements can be changed, to produce new heirachies (and therefore new images).
Impressive! Thanks for the release.
It's either that, or else there are multiple totally-unrelated methods of achieving essentially the same outcome.
Both might not do the computation using sequential IEEE 754 floating point operations, but they perform the computation nonetheless.
So this might be just a philosophical matter, but as far as I am concerned, there is no distinction between "performing a matrix multiplication" and "performing the operation that a matrix multiplication performs", such as projection, mapping, rotation, scaling. If you can mentally picture a plane being stretched, skewed, or spun, along with all the points on that plane, then you are mentally executing (in that case) 2x2 matrix operations. But we do this so effortlessly that we don't notice it.
As a kid learning to read, you might start by recognizing individual letters and sounding out the words. But at some point in fluency, you hardly even notice the letters anymore -- recognizing them disappears into an effortless task that is executed so quickly you might not even notice it's happening. That doesn't mean that your nervous system is no longer performing some kind of best-fit-comparison between an optical image on your retina and a set of learned characters; it certainly must be, but it's just happening (basically) with hardware acceleration.
You might think you aren't doing matrix multiplication because you aren't consciously iterating through [A1B1 + A2B1 + ...] like a child sounding out vowels, but the operation is necessarily happening somewhere along the line; it's just happening at the hardware level below your conscious perception.
I can do digital matrix multiplication with Numpy, but I can also make a circuit that does analog matrix multiplication using op-amps, where the addition and multiplication happens in voltages; I can make a mechanical device that does matrix multiplication using gears and levers, and I can certainly coax neurons to multiply matrices. It's hard to imagine a better structure to implement a matrix multiplication than the dendrites of a neuron.
Now I can just go run that on my computer.
insane.