Infinite Images and the Latent Camera
mirror.xyz
mirror.xyz
The old/existing way is: humans produce content (writing, images, videos, music) using software and then use the internet to host and distribute the content. As the volume of information increased, aggregators of content became more important. Google helps you find anything, but primarily text and images, YouTube indexes and recommends medium length videos, etc.
A very common problem pattern is that you find multiple sources that are _nearly_ what you want, but not quite. For example you combine two stock photos in photoshop (with some editing) to produce a final image.
The future way is: large foundational models that act as the aggregators of information _and_ the editing tools. Co-pilot is the best current example we have. Once the performance of Co-pilot becomes good enough, I will never have to leave VSCode + Co-pilot, replacing the previous pattern of Google + documentation + Stackoverflow. The indexers and aggregators will be made obsolete. This is exacerbated by the fact that Co-pilot is able to do more than just aggregating the correct Stackoverflow examples. Co-pilot can do the customization and translation of the code examples into your specific use case.
These large/foundational models are nowhere near perfect yet. Once the performance is good enough to make existing workflows easier, we'll see rapid adoption in every domain. Recommendation algorithms will look entirely different in 10 years time. The 2030 version of YouTube won't be serving users the closest match video from its library, it will be generating custom edits on-the-fly tailored to the user.
Collaborating with an external latent space and teasing certain viewpoints out of it is an interesting dance, and something that I really look forward to participating in as soon as I can get access to the tools.
I'm very curious about where the capacity for iterative refinement is/will be in the future as well. "Ok, that looks great, now can you try it with a bit more green?" etc.
[1]: https://i.redd.it/rfn3io4urcu81.jpg
"a painting of a bridge, giving me the satisfaction of having painted it myself"
No such thing appeared.I think that typing to create imagery is cool, but also agree that the ability to instantly create convincing images (with those images being considered the final work) is not the most exciting/rewarding part, at least when those images are isolated as singular works to consider(like in painting) and not part of a greater narrative. This perspective was part of our motivation to frame the subject differently!
Honestly, you could come up with some painting you would be proud of. But you'd actually have to engage with the tool and make something.
The verb paint is open to interpretation.
Creating stuff has multiple goals. I'm currently getting a lot of "satisfaction" from AI image generation. Does that rebut your point? It's hard to know...