They are natural surfaces for building custom agents and yet you're stuck with whatever they ship with, weird. It's not like it's too complicated api-wise either.
There must be something I ignore.
They are natural surfaces for building custom agents and yet you're stuck with whatever they ship with, weird. It's not like it's too complicated api-wise either.
There must be something I ignore.
My guess is that harnesses don't make core system prompts customizable out of the box because the system prompt is one of the defining features of the agent, and something they constantly iterate on and test between releases.
Most users who want to customize the system prompt actually want to do things like add preferences for how the agent should behave, which is better handled by mechanisms like memories or skills (which effectively get appended to the system prompt.)
Not only they get "lost" and ignored as the context grows, but the baseline behaviour of system prompts is retained in the agent.
Skills are prompts, albeit in a specific format. This is apparent in say, Codex where $MYSKILL is literally injecting the skill-prompt inline into a typed prompt. This all gets passed into the semantic memory system anyways, refining away cruft like redundancy, pleasantries, et al.
You have to remind it what's in it's own memory, or the subagent skill is influenced by the main system prompt.
It's sloppy vibe coders productivity porn.
> And none of this works properly. None.
`Properly` is an ambiguous term. It may mean "not what someone expects" or something else, but I think this assertion is factually incorrect.
> You have to remind it what's in it's own memory,
Almost every agent pushes out prompt/context-prompts from the context window (overloaded naming, fun), rather than treating the existing context-prompt hierarchy as an immutable part of the window. Regardless, looking at agent code, it works as designed.
> the subagent skill is influenced by the main system prompt.
This is implied and working as intended. The skill is an additive prompt. Prompts are part of a hierarchy. System, Agent, User, Input (skill and typed) which all influence each other. Why it's over-engineered with a format rather than a flat text file, is beyond me.
LLMs are virtually deterministic for small questions. They don't scale linearly or consistently. That's a mismatch of expectations, rather than a fatal flaw. The utility is a matter of risk management.
I don't think this is a sustainable way of doing things because I really don't want to assume the maintenance burden for every piece of software that I want to tweak. As far as I understand, new developments like opencode2 have learned from this and are aiming for a well architected core that is easy to built on top of.