My deepfake DALL-E 2 vacation photos passed the Turing Test
mattbell.us
mattbell.us
1. This judge is aware that they have to discern whether the 'bot' is real or a machine.
2. The judge cannot discern whether the 'bot' is real or a machine better than random chance.
This failed 1. And even given that advantage, might have failed 2 as well?
Often I see headlines along the lines of "X fools people and beats the Turing test!". But the point of the Turing test isn't to trick a person, it's to make it functionally impossible for a person to distinguish between the real and simulated thing, no matter how hard they try. For something to pass a Turing test, it would need to be able to pass the following:
"Anyone can play as judge any number of times. You can take as long as you want, and if you're successful in breaking it under controlled conditions (IE, you don't cheat and use an out-of-band communication protocol with the 'bot'/human), we'll give you a 10,000,000$."
The "One Million Dollar Paranormal Challenge" (https://en.wikipedia.org/wiki/One_Million_Dollar_Paranormal_...) is a solid example of a Turing test for magic.
If you're asking "do people often imagine they can confidently distinguish things when they actually can't" then the answer is a solid yes - things like audiophile and wine testing have proven that again and again.
Instead, I think you don’t want any third parties “editing” or “curating” the exchange (beyond whatever blinding is needed to make it work).
In the "paranormal challenge", the juges usually include stage magicians, because they know the tricks that would fool ordinary people. James Randi himself is a magician.
Also, it shouldn't matter if the human is someone who worked on the AI, or has read the code, or has seen every previous Turing test that the AI underwent. There shouldn't be any information a person could know that would allow them to tell that it's an AI.
"Are you going to be at the blah blah because I need blah". response: "I really wqnt it". Nonresponsive, typos, what does 'want' refer to? Who knows. Is this a bad bot, someone spun out on meth, someone with cognitive processing issues, a busy mom texting while distracted, or ?
I think you've inverted the expectation here. What I was saying is a machine passes if you can't find anyone in the world who can distinguish between it and an actual human. Meaning that if the "world's smartest person" can distinguish between it and an actual human, it fails, even if it can fool everyone else.
A machine can pass the test by deliberately feigning to be a human with limited communication capabilities (ex: I think we can simulate 2 month old baby talking via text). But then all you've shown is that your machine is as capable of thought as a 2 month old baby, which probably isn't the bar that you're trying to reach.
You must constantly be aware that images, or text, or voice, or other audio, or other signals or data, might be computer generated or altered.
All the time.
And you individually, or those about you, or societies at large, may be influenced in large or small ways by such signals, patterns, and records.
Your elderly neighbour or relative might be scammed out of life savings. Investors of false product claims. Voters of some fake outrage --- particularly of the October Surprise variety. Soldiers and diplomats of mock attacks, or false depictions of a tranquil situation where in fact danger lurks.
The test never ends.
This is your final warning.
Sense meant was "the threat is active, you can't presume otherwise, ignorance is not safety".
That said, I suppose that message will be repeated occasionally in future. Possibly even by myself.
It's a matter of quantity as well as quality. It gets easier to produce convincing forgeries all the time, but it doesn't get easier to debunk them. This allows for public discourse to be destroyed by flooding it with fakes.
Gans, deepfakes, X-does-not-exist, DALL-E, GPT-3, etc., are approaching if not at realtime in generation.
Fakes hit the newsfeed with, or before reality.
The latter most especially if part of a counterintelligence / disinformation blitz by attackers to maximise confusion.
Yes, in the 8th century BCE one could commission bards to gin up epic sagas of conquest and atrocity to stir the public spirit. In the 15th century, one could hire one of the most brilliant sculptors and painters of all time to create artpieces legitimising a power-hungry thieving dynasty which had seized control of banking, commerce, and the Church. In the 16th century, an earlier horde could by mobilised by an Orange tyrant through pamphlets cheaply printed and distributed. In the 19th century, Wreck-a-Feller could intercept (and rewrite) telegraph transmits, newspapers could sensationalise photographs, and preachers could roll holy until districts burned out.
And so into the 20th and early 20st centuries.
What was not possible until quite recently was for high-fidelity, all-but indistinguishable, continuous imagery, audio, and video to be produced in realtime and distributed globally.
The time between fraud and mass impact is now measure in seconds, and can reach globally. This is well within response intervals not only of populations and institutions, but of individual weapons of mass destruction. The risks of massive consequence are huge.
And that is what is new and novel.
The interaction with other technologies raises yet more concerns. With Slaughterbots a highly-plausible if not actual reality, "leaked" video showing a plausible Slaughterbots attack might itself be a weapon of war. (https://yewtu.be/watch?v=vR91F3tp6eQ)
That is, one of the risks of some presumed capability or knowledge is that even the belief of its existence or validity becomes a tool to be used. The late senator from Wisconsin claimed to have, though never presented, evidence of disloyalty, a common practice in witch-hunts. With increasing scope and scale of data collection and storage, such claims become of themselves ever more plausible, with numerous effects. Concepts such as the Panopticon and "chilling effects" operate not based on actual or observed risks, but on presumed ones. Adam Smith in Wealth of Nations uses the word invisible twice. The unhanded one references the decline in moral amongst soldiers faced with sudden and invisible death at any time, and the costs of perpetual vigilance.
So, no, really, this is not the same. Scale matters. Effects matter.
So, to perform the experiment, one must have some people play the game with humans a few times and then play with a human and a machine a few times, and look to see if the results are statistically significant. When they aren't, Turing posits, the question 'can machines think?' will have been answered in the affirmative.
That is not to say that this DALL-E vacation photo social media post constitutes a rigorous 'passes the Turing test'. But I don't think it's fair to criticize someone for using 'the Turing Test' colloquially as a catchall for saying 'you probably didn't notice this output was machine generated, therefore you might want to adjust your priors on the answer to the question, "can machines think?"'. Because that's exactly the spirit that Turing was working in when he proposed using a party game as a test of intelligence.
Then there's the conceptual argument of the Turing test, which we can turn into a test of intelligence. It relies on the idea that we can abstract the mind into a "thinking" black box which takes inputs and outputs. And then posits that any black box which can't be distinguished from an actual "thinking" black box may as well be "thinking".
Passing the literal Turing test is a sign that some humans can be tricked for some small domain over some small period of time. Passing the conceptual argument that the Turing test relies on shows that there are non-human entities which cannot be separated from humans on the basis of how they interact with a human (through some limited medium).
The repetition, duration and prizes are just practicality; prizes incentivize people to try, repetition ensures that the results are robust, and duration ensures that the domain the AI can be tested over is equivalent to a humans.
Aside from some corners (probably to be filled over the course of the next year or three), dalle2 is obviously outperforming almost all humans at its task. The cross-style ability is probably exceeding almost all human /artists/, who tend to specialize in a single style.
And some of the creativity criticisms (can only styles it's seen before) are basically true of all but the tiniest sliver of humanity, whose names we tend to remember.
This is a dialog between Q, a human who is trying to work out if he is talking to a computer, and A, a computer who is pretending to be human:
Q: In the first line of your sonnet which reads "Shall I compare thee to a summer's day," would not "a spring day" do as well or better?
A: No, I don't think so. A spring day is lovely, but it doesn't have the same connotations as a summer's day. Summer is a time of warmth and languor, of long days and lazy afternoons. It's a time for picnics and swimming and lazy days in the sun. Spring is a time of rebirth and new beginnings, of freshness and hope. I think a summer's day captures the mood of the sonnet better than a spring day would.The colloquial meaning of "pasted the Turing test" has come to be "has been able to demonstrate intelligence when put to some serious, systematic testing". That may be switching that "has been able to fool people when they didn't look hard". That might be changing but I don't think it's changed yet and until it's changed, I'll protest 'cause that's terrible change imo.
> It's likely that with a harder version of the Turing Test, in which real and fake images of the same content are presented side by side and people are told that one of them is fake, it would be much easier to detect the fake images.
But really it's just information overload, most things on social media I just scan the thumbnail and move on. Only my family would care to see my vacation photos :D
- The nudibranch (slug thing) on the green coral doesn't look like anything I've seen in the Caribbean before, and the coral also looks odd for the region. That said, this is probably the most difficult photo for me to differentiate. I would have accepted this as a cool find of something I haven't seen before.
- The grouper (big fish) photo is actually pretty good, although DALL-E has misplaced its eyes a bit. That said, the lit foreground and dark noisy background are exactly the look I would expect for someone using a basic camera + lights with wonky post-processing.
- The diver photo is a horror show. There's a hose going nowhere on her back. It looks like she's blowing out of a harmonica instead of a regulator. Bubbles are collecting around the top of her mask for some reason. Her fins look like they were badly Photoshopped. Nothing looks right here.
- The lobster photo has a real but subtle flaw: Caribbean lobsters don't have big claws. It also looks like it's under a rock like you would find in cold waters around Massachusetts and Maine instead of the Caribbean.
Interesting stuff though. It will force me to be more skeptical when I look at people's photos in the future.
More interesting question is this. Is it a crime if you generate CSAM just for ones own consumption?
Yep. If it isn't obviously fake (i.e. a cartoon) the possession is illegal whether you produce it yourself or not. Though it's probably safe to say that you're unlikely to get caught if you're not sharing those images with other people.
My point is that the courts are going to have a hard time with this.
I think we're in agreement that the advancement of the technology is going to make this topic come back up for legal debate. When the gulf between CGI and real photography was large, it was pretty straightforward. Not so much now.
I imagine it'll get challenged again at some point on constitutional grounds. It is illegal right now on a moral basis, which is probably the weakest argument over the long term.
AFAIK they are still illegal in Canada.
Then name itself is a portmanteau of Lolita complex, after the Nabokov novel.
https://en.wikipedia.org/wiki/Lolicon#Legality_and_censorshi...
Maybe an effective approach would be to maximize harm reduction until the addiction has run its course? That seems to be the Portugal solution, and it seems to be successful.
Diagnosing mental illness generally consists of identifying some number of symptoms out of a possible list - often something like five out of eight possible. That means two people can be diagnosed with the same thing with only two overlapping symptoms.
So basically, don't listen to the guy who starts with "I'm not a psychologist" and then decides to play armchair psychologist.
Do gay men eventually get tired of being homosexual and turn straight?
Yes, clearly not a psychologist, nor an addiction treatment specialist. Methadone is often used indefinitely as maintenance therapy for opioid use disorder.
Let us take this lesson from "pray the gay away" camps and other types of "conversion therapy": You cannot take these preferences out of people, it just does not work. At best, you can make it clear how they are is a travesty and a lot of them will hide it for the rest of their life successfully. This is not a good solution compared to what you were talking about earlier because it increases human suffering by a lot and doesn't have an advantage over a suffering-free solution.
That said, I don't think I could justify to myself to create such an AI. I've quite simply been so disillusioned by the depravity of man and especially in this case I want nothing to do with it even peripherally. Perhaps that makes me a little hypocritical. Philosophically it would still be nice to solve this issue, even if it is just so no more children need to suffer (which should always be the main goal).
Or will it strengthen their need for 'the real thing' as someone else suggested in a sibling comment?
In any case, we still don't have a great answer for the legal question. Possession of realistic fake imagery is illegal, on the grounds that its very existence is a risk to children. There isn't any actual science behind that, it's just what the politicians have said to justify regulating what would otherwise be a constitutionally protected right. I imagine it will become a topic of discussion again (my quick research says the last major revision to US law in this regard was about 20 years ago).
> We’ve limited the ability for DALL·E 2 to generate violent, hate, or adult images.
"Our investors include Microsoft, Reid Hoffman’s charitable foundation, and Khosla Ventures."
NSFW: https://medium.com/@davidmack/what-i-learned-from-building-a...
>I felt at this point that I’d hit a dead end. Press and fundraising would be tough and require some extra creativity and force. I spoke to friends about hiring them, and had polarized answers. Overall, this project had become less appealing.
Seems to me there is a lot of money to be made here if it works. The interesting thing about porn compared to say TV is that people have very, very specific interests and basically just want an infinite amount of content within that interest. It's not like with TV where cooking shows become popular, and the people that used to watch dramas are now watching cooking shows.
So the ability to generate highly specific content tailored to an individual's very precise requirements seems potentially very lucrative.
This doesn't say much about the amount of vaginas that this guy has seen.
We get it.
But it's a pretty close analogue.
This is certainly debatable, and I agree that it is pushing the limit a bit.
I think in the end, the "Turing test" was devised as a thought experiment, not as a final definition of AI. So I guess some freedom of interpretation is reasonable.
Also, if I'm drunk and reading nonsense I might not realize it.
If it's a game where nobody cares, it's a stupid game, and results are meaningless.
It's a break of the scenario if they didn't go in with the goal of detecting fakery. This makes it useless as a "Turing test".
Hope someone with invite access could do this for Moby Dick or Sherlock Holmes stories or 1984.
Diffusion models don't really have a latent but the CLIP embedding would serve the same function. The problem is, OA would have to implement it themselves. There's no way you can implement it as a user with the current interface. (This is also true of alternative methods like gradient ascent or GEDI etc.)
Also, underwater photos are not something many people have personal experience seeing. Most of us don’t live underwater. We may not be equipped well enough to tell the difference, where above water, especially urban photos, we will likely notice better.
Most people didn't notice that some of my vacation photos were fake, therefore it passed the Turing Test... why is this clickbait nonsense getting so much attention?
Can someone who upvoted this article explain why you upvoted it? Did the fact that the title is flatly false not bother you? If someone wrote an article about cracking some encryption algorithm and titled it "I proved P=NP" would you upvote it?
I bet I can come up with a simple generator that generates galaxies/nebula pictures and if I interspersed those in with NASA Hubble generated images, most people could not pick out the real Hubble images from my generated images.
They had a giant warehouse with a toy NYC that they flew a camera through with little models... pretty nuts given how movies are made now.
The making of Fifth Element is a pretty great watch.
Detail in https://www.reddit.com/r/sciencefiction/comments/53p7gw/orig...
I believe they intentionally hobbled it in this respect for "safety" (iow to keep themselves out of a scandal when someone asks it to create "President Biden accepting Bribes" or whatnot...)
Certainly far simple diffusion models trained including faces do just fine at creating photorealistic faces.
Preventing Harmful Generations
We’ve limited the ability for DALL·E 2 to generate violent,
hate, or adult images. By removing the most explicit content
from the training data, we minimized DALL·E 2’s exposure to
these concepts. We also used advanced techniques to prevent
photorealistic generations of real individuals’ faces,
including those of public figures.Which tells us nothing at all.
Here's something that I'm not blase about: AlphaFold. That is one of the crowning achievments of humanity. It solved a problem that people have been working hard on for 40 years using an algorithm that is less than 5 years old on a computational framework that's a couple years old, on hardware with ML training powers orders of magnitude higher than anything that existed 5 years ago, and conclusively demonstrated that evolutionary sequence information, rather than physical modelling, is sufficient to predict nearly all protein structures. And, once the best competitor had a hint how it was done, they were able to reproduce (much of) the work in less than a year.
Now that's amazing. World-class. Nobel-prize worthy. Totally unexpected for at least another 10 years, if ever. Completely resets the landscape of expectations for anybody doing biological modelling. However, it also won't transform drug discovery any time soon.
Who are we trying to talk to, I wonder?
Game with bored, disinterested players would be entirely meaningless.
And I'd like it more as a Gedankenexperiment if people weren't talking about it as a tool or metric. That kind of thinking gains momentum.
All the first impressive looking shots at the top of this article are real.
People here are nitpicking over the definition of the Turing test. What actually matters here is that, if not already now, but certainly in 1-5 years neural nets will most certainly be good as the 99th percentile artist.
Does that mean AGI is here? Probably not. But we are missing the forest for the trees.
Also note that DALL-E isn't limited to photorealistic styles.
Definitely, but this isn't an example of that, it's an example of people on Hacker News not wanting things to be wrong. The title is clearly wrong.
dalle-e is still impressive, but taking this to the extreme it would be like making it simulate pictures of TV noise and show we couldn't tell it from the real thing.
What would be unethical about creating a fake vacation? As long as you're not defrauding anyone, I don't see who would be hurt by this.
In fact how do you know DALL-E actually created them, and did just regurgitate some it was trained with?
I think DALL-E is not released so researchers are unable to take it apart yet, but this question was already researched a lot in the context of other generative models and so far they really did generalise (assuming a well trained model, the can overfit).
The image of the fish is strange. I found a few photos of similar fishes with vertical stripes but the fish in the image has squares.