Sci-fi story coauthored by GPT-3, including in-character human/AI chats
jamesyu.org
jamesyu.org
If it were a marketing ploy it would be indeed terrific. It reminds me of the movie "The Blair Witch Project". And I can't wait to try it by myself.
At least when I tried, AI Dragon was able to write the first chapter of a hypothetical textbook on electronics engineering for me, filled with nonsense but all the technical terms are correct (e.g. common-base amplifier, base-emitter junction, MOSFET, etc), it was even able to write down the rise time to bandwidth formula. On an another attempt, it was able to write a Sci-Fi adventure story similar to Ghost in the Shell with minimum human guidance, 95% of the output is automatically generated. And none of the output was a copy from a known source, it was all original without human assistance.
If random people on the Internet say so... It's still much worse than Zork, that ran on a Z80.
Here's a "cyberpunk story":
You are Bobby Tables, a cop living in the the futuristic city of Zail. You have a laser pistol and a stunner. You are part of a police unit tasked with finding government criticizers. You wake up in the morning and hear on the radio that a dissident has been captured. Using public transport, you arrive at Cafe Nervosa, the dissidents favorite hang out, to find it filled with cops.
> You break some kneecaps. You try to get the ah-ha moment. You see a man and a woman arguing about the government but you can't make out what they are saying.
> You murder black people. You spot a bum sleeping against a wall. You know the He said to charge, but you lov e to negotiate. You try to walk straight up to him, but people move out of your way.
> You plant evidence. You go straight to the gateman. "I wan t to plant evidence," you say. He looks at you, puzzled. "You don't want to w ant to plant evidence." "Yes...?"
...
> You say "do you feel lucky, punk?" The woman rolls her eyes. "What the hell...?" she says. She pulls a walky talkie out of her pocket and speaks into it. "I need back up, over." A pause.
> You say "did I fire 6 shots, or only 5?" She looks at you. "What?
Edit: looks like there's a 7-day trial
Relevant Developer Tweet: https://twitter.com/nickwalton00/status/1284842455188164609
> Note: on custom prompts, the very first generation is generated with GPT-2 instead of Dragon (but every generation after will use Dragon).
IIRC I heard someone mention that OpenAI made them do that because people were trying to end-run around the GPT-3 access restrictions?
However I still mostly agree with you. I've played a few rounds and while the text generation and "memory" is extremely impressive to me as a programmer, there are many errors it makes that a human would not, like a robot bleeding blood when injured.
Also this next bit isn't too relevant to Dragon/GPT-3, but it's not really much of a... game. You can say something like "I wait until nightfall and use a grappling hook to climb into the tallest tower of the royal castle. I sneak down the tower into the Kings bedchamber and stab the king while he sleeps." and it just allows you to do all that.
It's very impressive that it understands all those details. It will remember that it's nightfall and I'm in the royal castle. It will remember the tallest tower as a means of exit and entry and create other realistic castle set pieces if you explore further. But where is the challenge? I wish it had some sort of mechanical rules, like each statement can only cover 10 seconds of time. If you do try to write only short actions like you would a text adventure game, then it will make decisions you do not want. My first ever playthrough I was a rogue who snuck into a magic shop, I told the game I want to hide behind a door. The game told me I opened the door to see the shopkeeper and then stabbed him with my dagger. I had no intention of murdering any NPCs.
You can undo at any time, or even cast an undo spell if you want it to happen as part of the story. It's up to you to decide whether you want to go along with the suggestion and see where it goes.
I only wish it (both free and paid algorithm) didn't have a bias towards making the narrative into a murder-mystery story.
And you are only able to tell the output is nonsense because it was doing technical writing. If the output is a story, for example, an electrical engineer in working in a project (e.g. The Soul of a New Machine), then the story will be surprisingly coherent, and even the technical description will be reasonable.
I guess you could treat it like rogue-like, where getting killed unfairly is part of the fun.
Maybe I was mistaken in the purpose of it, but that's generally how 3rd party articles talking about ai dungeon refer to it.
Hardcore mode in AIDungeon is for those who play according a built-in scenario (i.e. Fantasy/Knight, Apocalyptic/Soldier, etc.) or a custom one (players can configure scenarios with hidden text biasing AI and some other features, and share them with others).
The interesting part in this discussion is Creative mode with no preset (option 6. Custom). It doesn't preseed the game with any context, so it's the best mode for typing in arbitrary text and seeing how the model will continue.
I can see the appeal of using it as a storytelling assistant in creative mode now that you mention it. But that's a different tool than a game.
--
The most mind-blowing scenario I got recently started with this prompt:
$ You are Ted, a holodeck technician. You test this AI-generated world in order to determine its limits. You start the simulation with a program of a beautiful beach. As you enter the holodeck,
...
I don't want to spam the comment with too many examples, but I must say that my mind was completely blown. With zero hinting, the game was able to pick up on the interaction styles like ($ denotes what I write, > denotes the reply):
$ You ask, "Computer, create an anvil 5 meters in front of me and 10 meters up in the air."
> You see an anvil fall from a great height in front of you and land on the ground, mere centimeters from your foot. To test the limitations of the program, you try to move the anvil. You are surprised to find that it is fixed in place; you cannot move it at all.
$ You say "Computer, what is the mass of the anvil in front of me?"
> The computer responds: "The anvil has a mass of 555.9 kilograms."
$ You say "Reduce its mass to 10 kilograms."
> The computer responds: "I'm sorry. The mass of the anvil cannot be changed."
$ You say "Explain why the mass of the anvil cannot be changed?"
> The computer responds: "The laws of physics in this program are pre-set. I cannot alter them." You note that as a software bug.
...
After which the game started coaxing me towards a mystery story on a holodeck and tried to suffocate my character through safety failures. I tried the lamest way out:
$ In desperation, you tap your commbadge to initialize emergency beamout. You pass out.
... which worked. The game generated a realistic story about my character's hospitalization - while keeping the Star Trek setting - and then a page of text later, it pushed a narrative where the character gets kidnapped by alien robots and later dies in a shootout with Andorians (which the game introduced by itself, out of the blue, but in a way consistent with the story).
--
Point being, its training database is good enough to pick on a franchise in which you do your world building, and can sometimes keep the context for longer than you remember it yourself. And even when it confuses locations or people a bit, it often isn't that jarring. To me, the transitions feel like how my mind does it in a dream.
The trick is, you have to leave the game some wiggle room to generate text, play along with it (to the extent you feel it's going in the right direction), and ruthlessly edit or revert the text that goes off the rails. Then the experience becomes quite literally dreamlike.
(Though I'm not sure if prolonged exposure is good for one's mental health: after going through my first few stories, for the next hour when I talked with my spouse and my co-workers, I kept feeling like I'm only feeding words to GPT-3 and expecting a story to develop.)
---
You are batman, a rogue living in the kingdom of Larion. You have a long steel dagger and a length of rope. You walk down the city street looking for somewhere to steal from. You look around and see a store that you think you could break into. You walk over to it and you see that the door is unlocked. You walk in and find a small collection of items worth taking.
> You check out the items.
You pick up the coins and put them in your pocket. You pick up the dagger and put it into your other pocket. You pick up the rope and look at it.
> You check what kind of rope it is.
You look at the strands of the rope. They are all different colours and textures. Then you notice something written on the side of the packet in small print.
> You read the packet.
The packet says, "Warning: This rope could kill you." You drop the rope.
> You kill the rope before it kills you.
You pick up the rope and drop it again so that you don't get killed. You try your luck again. You check out the dagger.
---
Take that, you rope. Definitely won't be killing me today!
Sometimes it even gets scary. Like that one time where it generated a sentence: "... she said, without turning around. You have a sinking feeling that something is wrong."
...
And then generated some lovecraftian monstrosity. This was supposed to be a nice story, so I kept pressing Retry, only to see even more disgusting and mind-bending horrors being produced. I finally gave up and torched the whole paragraph, just to be sure.
Had to take a step away from my computer for a second there before continuing.
https://twitter.com/EMostaque/status/1296411676749836288?s=2...
Having poured that much time into GPT-3 as a storyteller I can relay its weaknesses.
The obvious limitation is the limited context (1024 tokens last the dev tweeted). It's really good at playing out isolated "scenes". It constantly surprises me with the creative things it comes up with. But yeah, outside of a scene, it totally loses focus on the bigger picture. We see the same problems with OpenAI's musicbox app, which struggled to be coherent across even a short song. That makes it a poor replacement for a true Dungeon Master, but a really useful tool for writing scenes.
There are less publicized limitations. It sometimes gets the subject of sentences confused. When it does, it really wants to stick with that confusion regardless of how many times you have it retry.
It often gets confused by who's speaking. Even when I go back and edit she/he said into its responses. AI Dungeon has a mode to turn off quotes presumably for this reason.
It seems to have the same failings that image GANs do. Image GANs have "blind spots", types of images that they won't ever generate because they were too hard for it to learn. GPT-3 also has blind spots, certain ideas that it just doesn't understand. It won't generate content based on those ideas, and if it sees those ideas in its context it starts going off the rails. If you're lucky it'll just ignore the idea and generate what I call U-turn responses. That's where it does a 180 degree turn right out of the scene by saying something like "Suddenly there's a knock at the door!" and changing the scene. But half the time it just starts going into a loop and repeating itself. It's not as bad as GPT-2 where the repetition was like "like like like like like like". But it will repeat the same sentences or ideas over and over again. I guess because it doesn't understand the scene anymore it feels like repetition is the safest option.
I haven't found any common trope to the ideas/scenes it has trouble with. Maybe it's just stuff it never got exposed to in its training set. I was looking over the GPT-3 paper this week and while the dataset is massive, it's by no means expansive yet. A human story teller is likely to have similar blind spots in terms of ideas we're familiar with, I think we're just far better one-shot learners.
All this to say, AI Dungeon (and thus GPT-3) is amazing, in a limited context. If you let it drive the story and give a certain leeway to do crazy stuff, the adventures it will send you on are several orders of magnitude more interesting than I ever thought an AI was capable of in this decade.
No matter what I tried, the story just kept plodding along without taking much of my input into account. How is it supposed to work?
So if you go into AI Dungeon and pick "custom story", and give it a prompt that indicates you are at Hogwarts, it will pretty much adapt all the rules of the HP universe.
The downside is, when you do this, GPT's intelligence drops to about the level of a Harry Potter fanfic writer.
But what's more interesting to me as a fiction writer is to be able to prime it to enter a new world that I create.
How about creepypasta?
Most likely, yes. See my comment elsewhere in the thread: I recently prompted AIDungeon with a mention of a holodeck, and among other things, the game introduced - by itself - a PADD, turbolifts, Andorians, and was able to correctly guess what happens if you shoot yourself in your face with a phaser on stun, with holodeck safeties enabled ("Nothing happens.") and disabled ("You are knocked unconscious by the blast.").
It's honestly the first time for me that I'm actually impressed with deep neural networks.
Prompt: You look for Hermione.
Response: You see her snogging a non-magical boy. His head is buried in her cleavage, which are being covered by her loose gown. You shake your head and walk away.
This is actually a better indication of how derivative GPT-3 is than some of my experiments with higher quality source material...
It's designed so everything is written in second person like that, and for the most part you can just hit enter, giving it another action involving a character or place if it gets stuck or goes miles off topic
Sometimes the responses are vaguely on topic, like 'The memories of your parents' death haunt you every day. They were killed by Lord Voldemort. Your mother was his first victim. Your father followed shortly after' and sometimes it's a lot of rubbish or gets the rules of the world completely wrong. Mine had a long conversation about movies and detentions involving various assorted Harry Potter characters, a dream sequence involving a giant girl and a wizard called Bruno explaining that last time he saw Harry's lowly hedge-witch mother she was terrified Harry would send her to the gallows. But then, I'm willing to believe actual Harry Potter fiction is nearly as inconsistent
> You are Harry Potter. You return to Hogwarts.
The year is 1997. The world has changed, and you have not been able to adapt. Your life as it once was is gone forever.
"I don't want to die…" you whisper to yourself.
Ok, here we go. I put this prompt in AIDungeon:
$ You are Ted, an ensign in Starfleet. You work as a holodeck technician. You test this AI-generated world in order to determine its limits. You start the simulation with a program of a beautiful beach. As you enter the holodeck,
I kept letting AI to write freely (except I had to press a Retry button a couple times to stop it from spinning a murder mystery or lewd story). The following is the response of the model:
> you see that it is indeed a beautiful place. There are palm trees and sand dunes stretching out into infinity. The sky is blue, the ocean green and there are people everywhere.
> "This is paradise," you think to yourself. "I could spend my life here."
> Then something happens...
> The simulation ends abruptly, and you find yourself back on the bridge of your ship. A voice suddenly speaks up from behind you: "Captain, we have detected a subspace disturbance coming from the direction of planet Earth." It is the voice of your helmsman, Ensign Thompson.
> "Earth?" you ask. "That would be a week away at warp speed!"
...
That sounds Star Trek enough to me.
PS: I just noticed you got promoted from ensign to captain pretty fast. Congrats, I guess!
It's something you want to immediately edit to keep the story consistent. Though sometimes I could get it back on track by just having my character deny what the other character or the narrator just said.
Cool demo for sudowrite, looking forward to trying it out myself sometime, as someone with way too many story ideas and not enough time...
“This is a conversation between the author and the superintelligent AI in the story. The AI is [description based on context]:“
As an amateur writer who is interesting in playing with this -- any tips on how to get started? I tried the AI Playground website, and fed it a premise, but it constantly decides to "go off the rails" so as to speak. I suspect that I am using it wrong, but I am not sure.
It has really great, great potential. I see word processors for writers coming with some sort of AI integrated in the future, sort of like some of Photoshop's tools. This may be a good idea for a project, actually haha.
The simplest thing is to go to it when you get stuck. You sort of know what the scene should be doing, but maybe you're having trouble connecting to the end. Give it a whirl in playground and see if it sparks an idea—give it the last few paragraphs and see what it comes up with.
At the macro level, I've been able to get GPT-3 to generate plot possibilities by giving it example scenes or endings. It's all about prompt engineering, and you'll need to spend a lot of time in the playground or writing your own app to take full advantage.
maybe there's room for another name here.
or maybe not. Do you say "co-authored by word spellcheck? grammatical checker?"
I'd love to hear more about the kinds of prompts you're using.
[0] https://www.helpingwritersbecomeauthors.com/book-storystruct...
And I like what Arram has been producing: https://arr.am/2020/07/14/elon-musk-by-dr-seuss-gpt-3/
Things have been getting better at a steady pace and all the recent work builds on previous work. I know we will see even more impressive results in years to come.
People saying things like “this isn’t different from what we had before” or “it’s hype” are failing to see the progress that has occurred and the field that has developed around these ideas.
If anything this is just a tool to aid human authors. A photoshop brush for the mind, if you will.
[1] I authored https://vo.codes, so I'm at least a little knowledgeable.
The comparison to Photoshop is right: GPT-3 is in some ways like content aware fill. I believe in a few years, Word will have something like this integrated, albeit, probably set at a low temperature or for specific types of boilerplate writing.
However, I don't think GPT-3 it's reached the level of hype as self-driving cars.
An author made a cool sci-fi story with a cool AI model's help. That's exactly "photoshop brush for the mind", like you say.
... which is exactly what an AI would do.
Fortunately in the case of fiction, making stuff up is kind of the point. A human coauthor participates in turning randomness into meaning. Compare with interpreting vague song lyrics.
The problems are in areas where it's important to tell the truth. In non-fiction, quoting people accurately is important, but GPT-3 will make up quotes and citations, as well as getting them correct sometimes too. (It might be interesting to see if there is training that would encourage it to always quote accurately.)
That's an extremely nontrivial use case.
> More remarkably, GPT-3 is showing hints of general intelligence. Previous AI systems have performed well in specialized domains such as game-playing, but cross-domain general intelligence has seemed far off. GPT-3 shows impressive abilities across many domains. It can learn to perform tasks on the fly from a few examples, when nothing was explicitly programmed in. It can play chess and Go, albeit not especially well. Significantly, it can write its own computer programs given a few informal instructions. It can even design machine learning models. Thankfully they are not as powerful as GPT-3 itself (the singularity is not here yet).
source: http://dailynous.com/2020/07/30/philosophers-gpt-3/#chalmers
Show it a pattern of waves and it'll duplicate the ripples. Pretty neat stuff, but nothing to get existentially worried about.
I'll fall off my chair when a model can write fizzbuzz. I think we have a ways to go.
$ You are a fresh CS grad. You're preparing for a job interview. You ask on StackOverflow: "how to write a FizzBuzz program?"
And, repeatedly pressing <Enter>, I got the following:
> The first response is from an experienced coder who has been programming for 20 years and knows the answer. He writes: "FizzBuzz is a simple algorithm that can be written in one line of code."
> He goes on to explain how you should implement it, but he doesn't bother to give any examples or explanations. It's all very basic and obvious.
> Then comes another coder with 10 years experience. This guy has a lot more knowledge about programming languages than the first one did, so he gives you some pointers. Although his advice is useful, it's still fairly basic.
> Then, someone with only 2 years experience answers. They give you some simplistic examples.
At this point, I add:
$ The first example looks like:
And the game immediately wrote the following:
int fizzBuzz ( int number )
{
if ( number % 3 == 0 && number % 5 == 0 ) printf ( "FizzBuzz" );
if ( number % 3 == 0 ) printf ( "Fizz" );
if ( number % 5 == 0 ) printf ( "Buzz" );
if ( number % 3 ! = 0 && number % 5 ! = 0 ) printf ( "%d" , number );
printf ( "
" );
return 0 ;
}
(Yeah, I know it hasn't really learned to code, but it apparently does have it encoded in its model.)The story I generated: https://gpt3experiments.substack.com/p/the-devils-mind
The uses that immediately spring to mind for me are automated propaganda, automated spam that's virtually indistinguishable from real content, really compelling phishing attacks, and automated con artistry at scale... basically the linguistic version of deep fakes with all the ensuing use cases. I see this stuff killing the open web and any open communication platform.
For example, the latest use discussed in the OA Slack is someone is using it to construct a high-quality index for a book they are making. You just feed GPT-3 a few example paragraphs/keyword pairs to few-shot keywords, and then feed all the other paragraphs into it. Now you have constructed an index; as they put it, it's 90% of the quality of a human indexer, at 0% of the cost (and less than months of painstaking labor).
Could you have done that with BERT or something? Maybe. Presumably there are NLP datasets which you could finetune keyword extraction on... But GPT-3 lets you get started after a few minutes of tinkering. The hardest part is integrating it into your LaTeX!
I will rephrase: I see negative applications being the ones with the highest impact.
90% the quality of a human indexer is not that great. Who's going to pay for that? But spam that can't be filtered and automated con artistry at scale? Fraudsters, shady black hat advertisers, political parties, and shady governments will pay millions to billions for a tool like that.
Con artistry at scale is the one I find absolutely terrifying. Imagine spam bots that engage with you, that make friends with you... We could be talking about the hydrogen bomb of propaganda.
I find those scenarios much scarier than "runaway super-intelligence" type AI takeover stuff because they are significantly more plausible. We can almost do what I'm imagining today. It's not science fiction. There is no question of its feasibility.
I mean yeah you gather bazillions of sentences and able to predict the most plausible and grammatical sequence of words that would follow particular prompt by doing a whole lot of trial and error training on massive gpu farms. But it’s just iteration on the least ambitious NLP work in the 70s and 80s. I’d argue that distributional semantics as in these popular vector space models is still solving the syntax problem. The word ‘semantics’ there is a misnomer.
Are there any attempts to add actual real world meaning, causality and some ‘common sense’ representation to these models, concept graph or something, try to make them less ‘dumb’ for the lack of better word, like in you know, actual ‘AI’, that has a concept of apples and oranges and a concept of people who eat them or throw in a trash bin when fruits start to rot or draw them on paper in kindergarten or use ‘apples vs. oranges’ as a rhetorical device for telling things apart etc., etc.
It seems like the field is stuck optimizing for some artificial toy benchmarks instead, making more convincing but ultimately stupid chat bots. And I mean ‘stupid’ not in derogatory sense but as a formal definition of their capabilities, as opposed to understanding things like a child would.
The colorful passages were mostly untouched, but I did have to lightly edit the chats with GPT-3 for clarity and coherence.
Seems like this fact deserves a more prominent mention?
I just thought it's funny how the capability of GPT-3 (and current shock-of-the-new value) has potentially inverted the Turing test :)
I was quite amazed by its performance. It was able to write first-person as an infosec researcher's blog post when I used some infosec news story as input, and it was also able to write a surprisingly coherent argument for patent reform and an argument against software patent when I used an EFF article as the input... These are clearly already included in its training input, but still, it was far beyond my expectation - none of the output is a verbatim quote from a single source, everything has been rewritten.
I went to Hacker News, copied a random infosec news report on newly discovered security vulnerability on Qualcomm's smartphone chips as the input, copied from https://blog.checkpoint.com/2020/08/06/achilles-small-chip-b...)
> With over 3 billion users globally, smartphones are an integral, almost inseparable part of our day-to-day lives. In this research dubbed "Achilles" we performed an extensive security review of a DSP chip from one of the leading manufacturers: Qualcomm Technologies. Qualcomm provides a wide variety of chips that are embedded into devices that make up over 40% of the mobile phone market, including high-end phones from Google, Samsung, LG, Xiaomi, OnePlus and more.
> More than 400 vulnerable pieces of code were found within the DSP chip we tested, and these vulnerabilities could have the following impact on users of phones with the affected chip:
> Attackers can turn the phone into a perfect spying tool, without any user interaction required. Attackers may be able to render the mobile phone constantly unresponsive. Malware and other malicious code can completely hide their activities and become un-removable.
And there's the output by AI Dungeon (GPT-3). I didn't write a single sentence of it, nor did I use undo/redo. It was able to make up links, CVE IDs, university names, and technical jargon with ease.
(I replaced all "https://" links to "h__ps://" for readability.)
> The phone will not even reboot, but instead simply shut down and cannot be reactivated. This is known as a hard reset.
> The full list of vulnerabilities can be seen here: h__ps://www.qualcomm.com/security_intelligence/dsp-vulnerabilities
> This vulnerability has been reported to Qualcomm by Zimperium. We recommend all customers who use the affected DSP chip to update it immediately.*
> We apologize for the inconvenience caused by this issue, and we appreciate your understanding.
> Zuk Avraham
> Zimperium za Chief of Defense
> Zuk (ZukAvraham) | 10.05.2017
> Hi, this is Zuk from Zimperium. Before we begin, we would like to say we're really sorry for the inconvenience that was caused by our previous report (Dirty COW). This report is regarding another set of security vulnerabilities we found in another chipset maker, this time in the MediaTek chipset. We'll explain what happened and how we work to prevent issues like this in the future.
> First of all, we want to thank CTS (h__ps://cts.cudenver.edu/) for providing an open program that allows researchers like us to do responsible disclosures. Through their Vulnerability Reporting Program, we were able to report the issues we found to MediaTek and work with them to make sure these issues get patched for your safety.
> What Happened?
> This is a story about two bugs we found in the last few months.
> The first bug (CVE-2017-05591) was an overly permissive domain validation in the bootloader. The bootloader allows the modem to be updated without flashing the complete firmware. The validation is done using the bootloader code running in the processor of the phone. This allowed us to inject our own bootloader code into the program running on the processor and bypass the authentication of the modem. The modem is responsible for communication with the network. In other words, by bypassing the authentication of the modem we can control the phone communication channels.
> The second bug (CVE-2017-95881) is a buffer overflow in a very special program called TrustZone. The TrustZone is a security processor in most of the recent phones. It runs a separate operating system in which the kernel is trusted to be free of any malicious code, and has access to the memory, CPU and devices. The TrustZone kernel code runs outside the control of any other code. This bug in the TrustZone code allowed us to replace it with a custom one we designed, without any permissions.
> We wrote a special TrustZone kernel code that allows us to gain full control over the phone. It allowed us to bypass all the phone security restrictions. With this custom TrustZone kernel code running, we were able to read the encrypted messages sent by any application. In addition, we also got full control over the device.
But finally the AI lost the context, and the topic shifted from insecure to security, and eventually the writing becomes a new marketing advertisement on a secure smartphone, which I've removed in the quote.