In other words, relying on censoring the CoT can risk the effect of making the CoT altogether useless.
Basically: https://www.anthropic.com/research/reasoning-models-dont-say...
As far as I know deepseek is one of the few where you have the full chain of thought. Openai/Anthropic/Google give you only a summary of the chain of thoughts.
This is better thought of as another form of context engineering. LLM's have no other short-term memory. Figuring out what belongs in the context is the whole ballgame.
(The paper talks about the risk of training on chain of thought, which changes the model, not monitoring it.)
All of those seem like very reasonable criteria that will naturally be satisfied absent careful design by model creators. We should expect latent deceptiveness in the same way we see reasoning laziness pop up quickly.
Incorrectness doesn't required intent to decieve. It's just being wrong
None of this requires the ML model to have any interiority.
The ML model needn’t really know what a person really is, etc. , as long as it behaves in ways that correspond to how something that did know these things would behave, and has the corresponding consequences.
If someone is role-playing as a madman in control of launching some missiles, and unbeknownst to them, their chat outputs are actually connected to the missile launch device (which uses the same interface/commands as the fictional character would use to control the fictional version of the device), then if the character decides to “launch the missiles”, it doesn’t matter whether there actually existed a real intent to launch the missiles, or just a fictional character “intending” to launch the missiles, the missiles still get launched.
Likewise, if Bob is role playing as a character Charles, and Bob thinks that on the other side of the chat, the “Alice” he is speaking to is actually someone else’s role play character, and the character Charles would want to deceive Alice to believe something (which Bob thinks that the other person would know that the claim Charles would make to be false, but the character would be fooled), but in fact Alice is an actual person who didn’t realize that this was a role play chatroom, and doesn’t know better than to believe “Charles”, the Alice may still be “deceived”, even though the real person Bob had no intent to deceive the real person Alice, it was just the fictional character Charles who “intended” to deceive Alice.
Then, remove Bob from the situation, replacing him with a computer. The computer doesn’t really have an intent to deceive Alice. But the fictional character Charles, well, it may still be that within the fiction, Charles intends to deceive Alice.
The result is the same.
If something outwardly acts the same way as it would if it were a person, then the external effects are the same as they would be if it were a person. (<- This is nearly a tautology.) This doesn't mean that it would be a person, but it does mean that the concerns one would have about what outward effects it may cause if it it was a person, still apply.
(Well, assuming an appropriate notion of "outwardly" I guess. If a person prays, and God responds to this prayer, for the purpose of the above, I'm going to count such prayer as part of how the person "outwardly acts" even if other people around them wouldn't be able to tell that they were praying. So, by "outwardly acts", I mean something that has causal influence on external things.)
If something lacks interiority and therefore has no true intent to deceive, what does it matter that there was no true deception, if me interacting with that thing still results in me having false beliefs that benefit some goal that the thing behaves as if it has, to the detriment of my own goals?
My point is the observation is not support for such vibes referenced in og-parent post. You can’t just rearrange bits and hope for a spark of life. It is not gonna “come live.” The rabbit-hole dwellers keep trying to make everyone believe that. It is both amusing and tiring how persistent they are at this. It is understandable considering the money they stand to gain from gullible people rushing to invest.
If Napoleon Bonaparte had been replaced with a p-zombie he would have been no less capable of conquering.