Claude didn't care about it. When I pointed that out, it was apologetic and that was it.
Claude didn't care about it. When I pointed that out, it was apologetic and that was it.
Edit: I just asked Claude how it would interpret that and it said it could either mean that would not answer from memory alone and only anchor claims into things it can check, or it would tightly relate its responses to the context that I had supplied.
If it chose the latter, I could see why it wouldn't always resort to search results
Clear and recognizable technical vocabulary for engineers, or a legal context, to mathematical logic, philosophy, certainly in ML, take your pick. I would think it's pretty familiar to everyone who speaks English and if not still clear with context clues
An eval of this is likely worthwhile;
Re: "Grounded in logic" and "Grounded in theory"
Ground and justify all of the responses with logic and theory and real observations from qualified experiments with citations.
Present a coherent argument borne of logical premises with extant sufficient proven evidence of support. Assess and critique the response given such criteria that all responses should be valid logical arguments, and revise before responding
https://en.wikipedia.org/wiki/Logical_positivism#Decline_and...
Because LLMs also run off vibes and the writing style of your text, another important issue with your prompt here is that it makes you sound like a stuffy dork, or perhaps a pro se litigant. They won't respond to this well because LLMs have feelings too.
https://www.anthropic.com/research/emotion-concepts-function
Just be normal! And have evals.
Once there are - or next month when there will be - better models, agents, and agent harnesses for this, do you think that then we should concisely specify what is required instead of doing evals for particular models?
So meta-analysis and requisite language are too high-order for existing models and agents, and it's currently necessary to apply such procedural controls outside of the prompt?
I am not up to date on philosophy of science, but the scientific method is certainly always subjective, or at least can't be successfully expressed in a formal system.
Here's a book you can read: https://metarationality.com
> Once there are - or next month when there will be - better models, agents, and agent harnesses for this, do you think that then we should concisely specify what is required instead of doing evals for particular models?
Hmm, not sure what you mean. "Evals" are another way of saying "regression tests", so they're useful when you want to change or compare any part of the system.
> and it's currently necessary to apply such procedural controls outside of the prompt?
In general I think you should try to move controls out of the prompt and into an external system, but the downside is that it costs more, so it's not always necessary.
Do you think it is wise to optimize prompts for specific models or agents when there is a new model every month?
So, to build something like Co-Scientist the controls should be in the agent? Or RLHF'd like other things when training the model?
How well are LLMs able to reason about their own behavior?
Put another way, this entire post is about getting AI to not bluff. How do you know that it’s not bluffing in its response to you?
1: “safety”
2: proprietary protectionism and
3: it’s probably rapidly changing enough to be hard to pin down.
I am a massive proponent of this changing, it’s very strongly holding back these models. Change 1: models need to have more “LLM behavior analysis” in their training data/weight. Change 2: models need to have very detailed DETERMINISTIC logs of each step they take, and be able to access those logs. Change 3: the tuning and tweaking that happens more frequently needs to be in a .md file that the model can access.
Change 4: with the other changes done, the models should now engage in several self analysis steps layered into its whole thinking chain. “How did I reach this conclusion, did this require any guessing, do the key facts have research support online, quick check for Claudisms or AIisms and common LLM issues, did my changes alter underlying things like libraries without integrating them etc. etc.”
It’s how we think and refine our ideas and plans, and the models should mimic that.
They are not available to anyone. Nobody knows how they work, including the models, Sam Altman, God, etc. They're emergent from the training process.
> Change 4: with the other changes done, the models should now engage in several self analysis steps layered into its whole thinking chain. “How did I reach this conclusion, did this require any guessing, do the key facts have research support online, quick check for Claudisms or AIisms and common LLM issues, did my changes alter underlying things like libraries without integrating them etc. etc.”
Remember inference costs per token. Do you want to pay for this every time?