How well are LLMs able to reason about their own behavior?
Put another way, this entire post is about getting AI to not bluff. How do you know that it’s not bluffing in its response to you?
How well are LLMs able to reason about their own behavior?
Put another way, this entire post is about getting AI to not bluff. How do you know that it’s not bluffing in its response to you?
1: “safety”
2: proprietary protectionism and
3: it’s probably rapidly changing enough to be hard to pin down.
I am a massive proponent of this changing, it’s very strongly holding back these models. Change 1: models need to have more “LLM behavior analysis” in their training data/weight. Change 2: models need to have very detailed DETERMINISTIC logs of each step they take, and be able to access those logs. Change 3: the tuning and tweaking that happens more frequently needs to be in a .md file that the model can access.
Change 4: with the other changes done, the models should now engage in several self analysis steps layered into its whole thinking chain. “How did I reach this conclusion, did this require any guessing, do the key facts have research support online, quick check for Claudisms or AIisms and common LLM issues, did my changes alter underlying things like libraries without integrating them etc. etc.”
It’s how we think and refine our ideas and plans, and the models should mimic that.
They are not available to anyone. Nobody knows how they work, including the models, Sam Altman, God, etc. They're emergent from the training process.
> Change 4: with the other changes done, the models should now engage in several self analysis steps layered into its whole thinking chain. “How did I reach this conclusion, did this require any guessing, do the key facts have research support online, quick check for Claudisms or AIisms and common LLM issues, did my changes alter underlying things like libraries without integrating them etc. etc.”
Remember inference costs per token. Do you want to pay for this every time?