StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery
github.com
github.com
Demo of global directions: https://twitter.com/minimaxir/status/1378766961937555457
i've been documenting this theme in a twitter thread here https://twitter.com/dmvaldman/status/1358916558857269250
Or for Zoom. The Surrogates movie comes to mind.
Right now, many AI systems can receive instructions through python (which, to me, look like unnatural language but can be spoken). Systems like CLIP and systems built around the GPT models can take in massaged English language prompts and return an AI generated output based on that.
I think we will asymptotically approach having our systems “fully understand” human language but I also think we’ve already arrived at your implied future of communicating with them through an unnatural, intermediate language. Isn’t that exactly what programming is for?
Setting that aside, the starting point is not that they aren’t (going to be in the future) capable enough to understand us. It’s the opposite.
They will be so far ahead of us that they will have to dumb things way down for us to barely follow along with what’s happening.
Of course as is wise on HN you do carefully plant some weasel words. “This kind” of AI being the most obvious escape hatch for the defense of your argument. But I assume people are interested in the bigger picture AI, not just a narrowly defined AI like this repo only, or this approach only, or this git hash of this branch of this repo only, etc.
I’ve been building a product on GPT-3 [2] using extensive prompt engineering. It’s a bit like programming, a bit like writing. It’s kind of like giving instructions to a child, but a child with essentially infinite memory and perfect recall. Some tasks work quite easily via commanding, while others need quite a bit of massaging to get coherent results, like construction of entire fictional scenes or documents that would be found in the real world, but where you’re just looking for one paragraph of the document as the output.
I do think that as these language models mature, prompt engineering will go by the wayside. With minimal training, you’ll be able to tell the AI precisely what to do.
[1] https://www.gwern.net/GPT-3 [2] https://www.sudowrite.com/