I was working on creating a next-n-actions predictor for one of our use cases and not paying much attention for a PoC. I was fairly happy with the progress for a few days, before actually reading the eval code and seeing that we leaked the final state in every eval.
It's nice to let claude run loose on porting from framework to framework (port my code from TRL to NemoRL to Tinker to VeRL) but looking at what it does in the intermediate steps makes me want to claw my eyes out. And getting it to adhere to our domain model (e.g. we have an SFTConfig and a .to_trl(), or a Row and a .to_harmony()) is impossible.
Most of the time my pushbacks are true improvements, but I've seen a couple of instances where the LLM was happy to downgrade their own good solution.
I've used this a bunch as a suffix to try to prevent that, works OK in most cases, but not always obviously, works better in the system/developer prompt if you have access to those. Seems I've used that about ~1000 times since 2025/08 when I started using codex (- transcription duplications, so maybe 1/2 of that?).
$ rg -a -o "Answer grounded in truth" ~/.codex/sessions | wc -l
1046Although this was more of an issue when using opusplan as Sonnet loves to do this. I only use Opus now if using Claude.
Indeed, it's easy to surface this by sending one model a "Review" of their proposal to another, then bounce them back and forward, ask which one is best and both models will almost always say something like "The other proposal/review is better", I'm guessing because somehow they think it comes from the human, and "human is always right" or something.
Dude, the fucking model is great for sure, but there is nothing behind the illusion. It doesn't know if something is right or wrong - simpler or harder to reason about etc
It's just generating text, in a coherent manner while following rhetoric processes as a solid attempt at logical thinking
Why is that so hard for people to grok?
Our industry (and society after) is beyond doomed with people seeing these self affirmations as anything like "insightful" validation.
That fundamentally wouldn't happen if it wasn't just an illusion.
There is value in it for sure and I can use it to write a lot of simple code, which is 99.99% of enterprise software - but that's another topic.
Something that can write a correct code snippet or even larger program that accepts the correct input and provides the correct output and otherwise is consistent with the given spec is doing something substantially more than just autocomplete.
> It's just generating text, in a coherent manner while following rhetoric processes as a solid attempt at logical thinking
So yeah, I do agree that they can make a very reasonable amount of reasoning. As a matter of fact, they reason about things better then an average Joe off the street ime.
That's entirely unrelated to what I said though, I think you misinterpreted/misunderstood what I wrote earlier.
They can make solid attempts at reasoning, its just not grounded in reality. It just applies these rhetoric processes to the current text - but it doesn't understand wherever it's actually correctly reasoned. Hence the answer "you're right to push back on this" is just the model being a sycophant. The sentence does not mean that anything of value has been communicated in either direction, and thinking that it has means the person in question is suffering from ai psychosis