I have a pretty large set of prompts that go into any software engineering, and I force every single agent to use an ephemeral style stack of prompt management. So, every turn it goes to the top of the stack and it is the very last thing they see in terms of all of my prompts and instructions and agent files. And then it gets taken out of the conversation so that it doesn't get sent to the agent the next turn (no context bloat). It has restored so much sanity.
I tried the caveman add-ons, and I felt like I was losing IQ points because I spend a lot of time reading agent output, and when they start talking like cavemen, I start thinking like cavemen. That was not good for my mental health. So, I try and make the agent talk like me and think like me. And it works, mostly. And my observation is that maybe I'm not the most efficient agentic thought process, but my sanity is retained.
All of that is to say that if something is reading like that to you, just have the agent rewrite it and read it in a rewritten tone because it's probably bad as it stands and your colleague did not put enough effort in it. It is /not/ good and you should not accept it as a default. We have to hold the line on stuff like this and maintain some semblence of normal human engineering standards that existed before AI. They are not making us better. They are making is lazy and dumber.
Opus 5 and other agents in the latest rounds of tuning have gotten ridiculously bad in terms of how they feel to interact with with all the invented language and localized nomenclature. It is an obvious bias that big words and technical talk looks good to the bottom of the bell curve, but when you actually try and understand it, it's horrible. So people say, "Yeah, that looks great," in all the RLHF rounds, and they run with it because they think it looks good, but it doesn't. It's terrible.
Hold the line. It isn't you. And it isn't a good methodology document.
Uh - dude - this means you're paying 10x in token costs because there's no caching.
If you 're-write token history' then you can't cache tokens.
It means for any reasonably long conversation, the llm has to reprocess the entire history as preflow on every prompt.
Are you sure you're really doing what you say you're dong, and how is it not blowing up your budget?
Perhaps more persnickety, it pushes the LLM out of distribution - if it’s unnatural for it to write in plain language without the prompt stack, your prefix will be an unnatural conversation which can reduce intelligence in hard to measure ways, especially over long conversations.
Not saying don’t do it, clarity is perhaps worth the intelligence hit, but it’s not going to be a free lunch.
I feel I fight them less with this setup. They get so lost in their own invented bullshit they stop being useful pretty often without it. So I would take bets on that :)
Did you have to build your own harness for this? Or hack Claude Code or something?
I thought it was maybe me just becoming lazier with reading considering how much AI-generated text I'm subjected to against my will, but I picked up Blood Meridian the other day and have unironically had less trouble parsing that novel than I have the majority of LLM-authored text I've seen.
Worse, I noticed that people in an office environment themselves have adopted a more speculative, communication style.
In the past, people remembered what was said and would draw attention to discrepancies. I could trust what people said.
Nowadays it's like; someone can say one thing one day and the opposite the next day (through convoluted language) and nobody bats an eyelash. Or sometimes someone will agree with me but then what they say immediately after reveals that they didn't understand the essence of my point at all. I didn't notice these things 5 years ago.
I guess this is what AI researchers refer to as 'model collapse' - it seems to affect people too though...
It feels like people don't value knowledge as they used to.
It's really hard to avoid mistakes when everyone is subtly covering them up. It feels like a lack of care and I find it demotivating.
I think because engineers are afraid for their job, they are under more pressure to talk a big game. Also under more pressure to deliver short term visible results. Bad combo.
If I had to, I'd process it into a short summary and/or ask an agent questions I have about the methodology.
I would then give my feedback. If they ask for details, I can have my agent update the document directly too.
I would not treat an AI generated artifact, be it documents or code, as something that a human should process fully manually.
I think that's where the future is going. It presents interesting challenges, but also some opportunities to make life better for everyone.
[1*] https://foundrs.com/have_your_ai_talk_to_my_ai.html
(*) I dislike the HN trend of starting citations at 0, I find it snobbish.
And when you ground it with real data, it's actually extremely useful. It's not exactly like Claude Therapist, but it's sort of the teach me about philosophy, but actually grounded and not vied. I have a lot of really strict prompts and grounding and agentic guidelines for this particular agent flow and harness that I've built.
And it's just a few weekends of vibing and feeding it basically all of Wikipedia and several gigabytes of papers and stuff, but it actually leads to interesting discussion. So I just have my personal philosophy bot and it's pretty fun.
One of the modalities I built is having two agents assume a famous persona. And then they take a thing, like grief or some thing that I experienced during the week, and they assume the role of the two different philosophers, and I just have them go back and forth 30, 40, 50 turns. And it's actually quite interesting, and it really moderates their language and tonality and behavior. They really get into the roles when you have the right prompting and grounding. Sometimes they get a little off the rails, but it leads to genuinely interesting areas to explore, and then I'll actually go read source material and things like that. I don't know, that's how I do therapy these days, but I never actually did therapy, so I just think a lot, now with agents finding interesting stuff to think about too!