At least when I tried, AI Dragon was able to write the first chapter of a hypothetical textbook on electronics engineering for me, filled with nonsense but all the technical terms are correct (e.g. common-base amplifier, base-emitter junction, MOSFET, etc), it was even able to write down the rise time to bandwidth formula. On an another attempt, it was able to write a Sci-Fi adventure story similar to Ghost in the Shell with minimum human guidance, 95% of the output is automatically generated. And none of the output was a copy from a known source, it was all original without human assistance.
If random people on the Internet say so... It's still much worse than Zork, that ran on a Z80.
Here's a "cyberpunk story":
You are Bobby Tables, a cop living in the the futuristic city of Zail. You have a laser pistol and a stunner. You are part of a police unit tasked with finding government criticizers. You wake up in the morning and hear on the radio that a dissident has been captured. Using public transport, you arrive at Cafe Nervosa, the dissidents favorite hang out, to find it filled with cops.
> You break some kneecaps. You try to get the ah-ha moment. You see a man and a woman arguing about the government but you can't make out what they are saying.
> You murder black people. You spot a bum sleeping against a wall. You know the He said to charge, but you lov e to negotiate. You try to walk straight up to him, but people move out of your way.
> You plant evidence. You go straight to the gateman. "I wan t to plant evidence," you say. He looks at you, puzzled. "You don't want to w ant to plant evidence." "Yes...?"
...
> You say "do you feel lucky, punk?" The woman rolls her eyes. "What the hell...?" she says. She pulls a walky talkie out of her pocket and speaks into it. "I need back up, over." A pause.
> You say "did I fire 6 shots, or only 5?" She looks at you. "What?
Edit: looks like there's a 7-day trial
However I still mostly agree with you. I've played a few rounds and while the text generation and "memory" is extremely impressive to me as a programmer, there are many errors it makes that a human would not, like a robot bleeding blood when injured.
Also this next bit isn't too relevant to Dragon/GPT-3, but it's not really much of a... game. You can say something like "I wait until nightfall and use a grappling hook to climb into the tallest tower of the royal castle. I sneak down the tower into the Kings bedchamber and stab the king while he sleeps." and it just allows you to do all that.
It's very impressive that it understands all those details. It will remember that it's nightfall and I'm in the royal castle. It will remember the tallest tower as a means of exit and entry and create other realistic castle set pieces if you explore further. But where is the challenge? I wish it had some sort of mechanical rules, like each statement can only cover 10 seconds of time. If you do try to write only short actions like you would a text adventure game, then it will make decisions you do not want. My first ever playthrough I was a rogue who snuck into a magic shop, I told the game I want to hide behind a door. The game told me I opened the door to see the shopkeeper and then stabbed him with my dagger. I had no intention of murdering any NPCs.
You can undo at any time, or even cast an undo spell if you want it to happen as part of the story. It's up to you to decide whether you want to go along with the suggestion and see where it goes.
I only wish it (both free and paid algorithm) didn't have a bias towards making the narrative into a murder-mystery story.
And you are only able to tell the output is nonsense because it was doing technical writing. If the output is a story, for example, an electrical engineer in working in a project (e.g. The Soul of a New Machine), then the story will be surprisingly coherent, and even the technical description will be reasonable.
Hardcore mode in AIDungeon is for those who play according a built-in scenario (i.e. Fantasy/Knight, Apocalyptic/Soldier, etc.) or a custom one (players can configure scenarios with hidden text biasing AI and some other features, and share them with others).
The interesting part in this discussion is Creative mode with no preset (option 6. Custom). It doesn't preseed the game with any context, so it's the best mode for typing in arbitrary text and seeing how the model will continue.
I can see the appeal of using it as a storytelling assistant in creative mode now that you mention it. But that's a different tool than a game.
I guess you could treat it like rogue-like, where getting killed unfairly is part of the fun.
Maybe I was mistaken in the purpose of it, but that's generally how 3rd party articles talking about ai dungeon refer to it.
--
The most mind-blowing scenario I got recently started with this prompt:
$ You are Ted, a holodeck technician. You test this AI-generated world in order to determine its limits. You start the simulation with a program of a beautiful beach. As you enter the holodeck,
...
I don't want to spam the comment with too many examples, but I must say that my mind was completely blown. With zero hinting, the game was able to pick up on the interaction styles like ($ denotes what I write, > denotes the reply):
$ You ask, "Computer, create an anvil 5 meters in front of me and 10 meters up in the air."
> You see an anvil fall from a great height in front of you and land on the ground, mere centimeters from your foot. To test the limitations of the program, you try to move the anvil. You are surprised to find that it is fixed in place; you cannot move it at all.
$ You say "Computer, what is the mass of the anvil in front of me?"
> The computer responds: "The anvil has a mass of 555.9 kilograms."
$ You say "Reduce its mass to 10 kilograms."
> The computer responds: "I'm sorry. The mass of the anvil cannot be changed."
$ You say "Explain why the mass of the anvil cannot be changed?"
> The computer responds: "The laws of physics in this program are pre-set. I cannot alter them." You note that as a software bug.
...
After which the game started coaxing me towards a mystery story on a holodeck and tried to suffocate my character through safety failures. I tried the lamest way out:
$ In desperation, you tap your commbadge to initialize emergency beamout. You pass out.
... which worked. The game generated a realistic story about my character's hospitalization - while keeping the Star Trek setting - and then a page of text later, it pushed a narrative where the character gets kidnapped by alien robots and later dies in a shootout with Andorians (which the game introduced by itself, out of the blue, but in a way consistent with the story).
--
Point being, its training database is good enough to pick on a franchise in which you do your world building, and can sometimes keep the context for longer than you remember it yourself. And even when it confuses locations or people a bit, it often isn't that jarring. To me, the transitions feel like how my mind does it in a dream.
---
You are batman, a rogue living in the kingdom of Larion. You have a long steel dagger and a length of rope. You walk down the city street looking for somewhere to steal from. You look around and see a store that you think you could break into. You walk over to it and you see that the door is unlocked. You walk in and find a small collection of items worth taking.
> You check out the items.
You pick up the coins and put them in your pocket. You pick up the dagger and put it into your other pocket. You pick up the rope and look at it.
> You check what kind of rope it is.
You look at the strands of the rope. They are all different colours and textures. Then you notice something written on the side of the packet in small print.
> You read the packet.
The packet says, "Warning: This rope could kill you." You drop the rope.
> You kill the rope before it kills you.
You pick up the rope and drop it again so that you don't get killed. You try your luck again. You check out the dagger.
---
Take that, you rope. Definitely won't be killing me today!
Sometimes it even gets scary. Like that one time where it generated a sentence: "... she said, without turning around. You have a sinking feeling that something is wrong."
...
And then generated some lovecraftian monstrosity. This was supposed to be a nice story, so I kept pressing Retry, only to see even more disgusting and mind-bending horrors being produced. I finally gave up and torched the whole paragraph, just to be sure.
Had to take a step away from my computer for a second there before continuing.
https://twitter.com/EMostaque/status/1296411676749836288?s=2...
The trick is, you have to leave the game some wiggle room to generate text, play along with it (to the extent you feel it's going in the right direction), and ruthlessly edit or revert the text that goes off the rails. Then the experience becomes quite literally dreamlike.
(Though I'm not sure if prolonged exposure is good for one's mental health: after going through my first few stories, for the next hour when I talked with my spouse and my co-workers, I kept feeling like I'm only feeding words to GPT-3 and expecting a story to develop.)
No matter what I tried, the story just kept plodding along without taking much of my input into account. How is it supposed to work?
Relevant Developer Tweet: https://twitter.com/nickwalton00/status/1284842455188164609
> Note: on custom prompts, the very first generation is generated with GPT-2 instead of Dragon (but every generation after will use Dragon).
IIRC I heard someone mention that OpenAI made them do that because people were trying to end-run around the GPT-3 access restrictions?
Having poured that much time into GPT-3 as a storyteller I can relay its weaknesses.
The obvious limitation is the limited context (1024 tokens last the dev tweeted). It's really good at playing out isolated "scenes". It constantly surprises me with the creative things it comes up with. But yeah, outside of a scene, it totally loses focus on the bigger picture. We see the same problems with OpenAI's musicbox app, which struggled to be coherent across even a short song. That makes it a poor replacement for a true Dungeon Master, but a really useful tool for writing scenes.
There are less publicized limitations. It sometimes gets the subject of sentences confused. When it does, it really wants to stick with that confusion regardless of how many times you have it retry.
It often gets confused by who's speaking. Even when I go back and edit she/he said into its responses. AI Dungeon has a mode to turn off quotes presumably for this reason.
It seems to have the same failings that image GANs do. Image GANs have "blind spots", types of images that they won't ever generate because they were too hard for it to learn. GPT-3 also has blind spots, certain ideas that it just doesn't understand. It won't generate content based on those ideas, and if it sees those ideas in its context it starts going off the rails. If you're lucky it'll just ignore the idea and generate what I call U-turn responses. That's where it does a 180 degree turn right out of the scene by saying something like "Suddenly there's a knock at the door!" and changing the scene. But half the time it just starts going into a loop and repeating itself. It's not as bad as GPT-2 where the repetition was like "like like like like like like". But it will repeat the same sentences or ideas over and over again. I guess because it doesn't understand the scene anymore it feels like repetition is the safest option.
I haven't found any common trope to the ideas/scenes it has trouble with. Maybe it's just stuff it never got exposed to in its training set. I was looking over the GPT-3 paper this week and while the dataset is massive, it's by no means expansive yet. A human story teller is likely to have similar blind spots in terms of ideas we're familiar with, I think we're just far better one-shot learners.
All this to say, AI Dungeon (and thus GPT-3) is amazing, in a limited context. If you let it drive the story and give a certain leeway to do crazy stuff, the adventures it will send you on are several orders of magnitude more interesting than I ever thought an AI was capable of in this decade.