While temperature can mix things up a little, we’ll quickly start to see how the depth of LLM output is tied to the depth of LLM input — and people’s prompts are not all that varied and deep at the end of the day.
We’ll start to notice signatures to each generation of model in its responses to tiringly common prompts — especially in stuff like blog spam, homework essays, email punchups, casual fiction, etc
There are certain poles that it gravitates to that end up reading like verbal tics or lazy ideas, and those will be more obvious as we begin to get inundated with generated content.
Careful prompt work can evade those poles and tics, but most people won’t have the skill or drive to bother with that.
I absolutely agree with this.
Furthermore, I've come to believe that "careful prompt work" doesn't really save much time relative to just writing/coding the output I want.
Especially after QA/edits
The real magic happens when you train a better AI using those perfected prompts.
temperature can actually go up to 2 in the API and you can also go use top_p < 1. but for the more out of this world stuff, this is where fine tuning will come in. The main model will always come in as vanilla anyway.
Something like: "We're gonna start a text RPG game. Every animal we meet should be a made-up animal composed of joining two existing ones, with creative names being made for them. Don't use any usual fantasy tropes. Make encounters have surprising but logical outcomes" and so on.
A human dungeon master doesn't pull a random enemy from the aether super well (or at least I don't!!). When I pick monsters I either logically think what makes sense for the area, pick something that's cool or vaguely humorous that I'm feeling the vibe of, or... use a random encounter table.
For treasure and rewards... if it's not something hand picked, I'm again getting the loot tables out and rolling. And heck even for gold rewards and things, I'm rolling the dice.
After trying basic stuff with langchain, an AI integrated with these types of tools is really promising. I think more systems on top could make it into a passable DM- at least until context becomes an issue.