LangChain Agent Simulation – Multi-Player Dungeons and Dragons
python.langchain.com
python.langchain.com
It has been a lot of fun working on it because it gave me a taste of new way of working, honing prompts and finding the limits of the AI, then figuring out how to offload the parts it struggled with (mainly storing stats, inventory, regurgitating room descriptions, etc.).
The “limits of the ai” for this kind of thing is basically everything other than pure dialogue in my experience.
I also had a bit of a play with this, but the LLMs were just rubbish at it.
The only meaningful way of doing this is to write an actual game using actual state, and use the LLM as a “renderer” that renders coherent state into free text, and free texts into (more or less) structured action requests.
Without rails, it just becomes like AI dungeon; free wheeling adhoc story telling with no rules or structure…
Large context windows don’t solve large scale coherence, and prompt engineering does sfa against the devoted trolling efforts of actual players.
That's all folks!
Step 1. Fine-tune a base LLM to the scenario. Feed it as much background material as possible. This would work best for a franchise with a huge extended universe or associated works: Dungeons & Dragons, Star Wars, Doctor Who, etc...
Step 2. Fine-tune / RLHF with negative weights against anything out-of-context. Basically, stop the AI ever referencing anything that can't exist in the fictional universe. Penalise references to real-world events, places, or people, modern technology, etc...
Step 3. Fork the AI model, once for each character. Fine-tune for conversations in that character's "tone" or mannerisms, backstory, etc... These can be via RLHF generated with a powerful model such as GPT 4. Again, reward/penalise the AI if it references anything it should or shouldn't know from that character's perspective.
When players play the game and converse with characters:
1. They'd always be talking to an LLM fine-tuned to death for that specific character in that fictional universe.
2. Then have both the input and output run past a general-purpose "nanny" AI that is prompted to look for exploits, out-of-context shenanigans, or unexpected output. Respond to the user with "I don't understand the strange things you're talking about" or some similar general push-back against jailbreaks.
Alternatively, have the generic security filter AI rewrite inappropriate terms in user inputs with unintelligible garbage. E.g.: if the user asks
"Are you a computer?"
Rewrite that to: "Are you a gizwallop?"
Then the in-game character would rightly be confused by the nonsense term and answer something like: "I have no idea what you mean, what is a gizwallop?"
Which could be translated back to: "I have no idea what you mean, what is a computer?"
This would be absurdly expensive and slow right now, but in 5 years? 10?I can imagine the cost of tuning the above for a GPT5-equivalent model dropping down to the budget of even a tiny indie game, let alone big-budget AAA game studios.
I joke but it is terribly jarring when the API is working perfectly and then starts apologizing that it cannot do something, like access personal information, when it is internally prompted that it should only use information it receives in the prompt.
Sure, the artificially applied constraint in the APIs also exist... and sure, you can have a long context with for example, gpt-3.5-turbo-16k, but the problem is fundamentally that no amount of wishing can make an LLM into a compiler that executes code.
You cannot, and will probably never be able to define your constraints in free text to an LLM of this type and then expect it to also be able to execute those constraints in an error free manner. That's not how the technology works. You might be able to make it generate procedural code that satisfy the constraints and execute that code in a reliable procedural manner, but afaik no one has managed to get that to work reliably and at scale (if you're thinking of smol developer right now, you clearly haven't actually used it).
When you define an RPG system to an LLM, the issue isn't that it isn't allowed to say things, it's just that it cant follow the rules reliably, and it can't keep track of what's going on as the context length gets larger and larger.
...and, for an RPG system, where the contrived RPG rules and internal consistency are everything, it's a deal breaker.
> when the API is working perfectly and then starts apologizing that it cannot do something
eh, use an open LLM. That's 100% not the problem.
https://blog.katarismo.com/2023-05-26-i-m-a-dad-i-replaced-m...
I'm doing all of it in JavaScript. I'm still not excited about langchain, and the JS version is lagging a lot from the python version. By the way, pocketbase is amazing for something like that.
I'm using JS to wrap around the llama.cpp project. Most of the work is done by that. JS is just used to pull the tokens out, normalize them, and then send them to a real-time database (pocketbase). The frontend is a Svelte app.
It's not a lot of code, llama.cpp does almost everything.
Not much here though, just rotating chat with the LLM taking on both sides. The prompts (https://python.langchain.com/docs/use_cases/agent_simulation...) are fine, but not great. No guidance on tension, on actions failing or succeeding, on plot advancement. Nothing to break out of anticipation loops, which are common (when the LLM promises something "is about to happen" but doesn't actually know what and keeps deferring the action).
It probably will work OK because the LLM plays both sides, and does so "fairly", i.e., always in character and never trying to "win". It'll run out of space for the history, but simple pagination might fix it (maybe that's even a LangChain feature?) – it'll drift, maybe dramatically. Or, given the system prompt, it might _not_ drift when appropriate; that is, it might not allow diverging from the original concept, and so not allow consequential action.
> I know that they are afraid of fire
The player side is doing a lot of heavy lifting here in actually directing the game.
I see two contenders:
https://github.com/minimaxir/simpleaichat/tree/main/simpleai...
https://github.com/griptape-ai/griptape
There is also the llm command line utility that has a very thin underlying library, but which might grow eventually: https://github.com/simonw/llm
https://github.com/lgrammel/modelfusion
It lets you stay in full control over the prompts and control flow while make a lot of things easier and more convenient.
import openai
import os
openai.api_key = os.environ.get('OPENAI_API_KEY')
def completion(messages):
response = openai.ChatCompletion.create(
model = gpt_model, temperature = 0, messages = messages )
return response['choices'][0]['message']['content'].strip()
response = completion([
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Who won the world series in 2020?"} ])
#####
import json
import tiktoken
import os
tokenizer = tiktoken.get_encoding("cl100k_base")
class Message:
def __init__(self, role, text, length=None):
self.role = role
self.text = text
if length != None:
self.length = length
else:
self.length = self._count_tokens(text)
print("New message, token length is",self.length)
def _count_tokens(self, text):
tokens = tokenizer.encode(text)
return len(tokens)
class History:
def __init__(self, ID=None):
self.messages = []
self.ID = ID
if self.ID:
self._load_from_json()
def add(self, role, text):
message = Message(role, text)
self.messages.append(message)
self._save_to_json()
def _save_to_json(self):
if not self.ID:
return
data = {
"messages": [{"role": m.role, "text": m.text, "length": m.length} for m in self.messages]
}
self.create_dir_if_not_exists('conversations')
with open(f"conversations/{self.ID}.json", "w") as f:
json.dump(data, f)
def create_dir_if_not_exists(self, directory_path):
if not os.path.exists(directory_path):
os.makedirs(directory_path)
def _load_from_json(self):
try:
self.create_dir_if_not_exists('conversations')
with open(f"conversations/{self.ID}.json", "r") as f:
data = json.load(f)
self.messages = [Message(m["role"], m["text"]) for m in data["messages"]]
except FileNotFoundError:
pass
def recent_messages(self, max_tokens):
recent_messages_reversed = []
total_tokens = 0
for m in reversed(self.messages):
if total_tokens + m.length <= max_tokens:
recent_messages_reversed.append({
"role": m.role,
"content": m.text
})
total_tokens += m.length
else:
break
recent_messages = recent_messages_reversed[::-1]
return recent_messages for m in reversed(self.messages):
if total_tokens + m.length <= max_tokens:
recent_messages_reversed.append({
"role": m.role,
"content": m.text
})
total_tokens += m.length
else:
break
It would be important to change that to not drop system prompts, ever. Otherwise a user can defeat the system prompt simply by providing enough user messages.Guidance (microsoft) - almost abandoned - https://github.com/microsoft/guidance