The reason "the agent suddenly started suggesting all kinds of things to make its code more robust" is because you said you "want to build reliable software".
It's not a signal of good judgment or understanding. It's just how LLM attention works.
EDIT: I mean, those systems accumulated so much complexity around the attention based next token predictor.
1. LLM thinking 2. RLHF 3. The latest frontier models
that does anything to change this fundamental "suggestibility" of LLMs.
But who knows, maybe I'm wrong.
this feels like "make no mistakes" level of prompting. reliable software isn't as simple as making it reliable, it's about choosing the trade-offs in the areas that don't matter as much as the areas that do. if you keep prompting the LLM to make your software more robust it will keep giving you things to do. they aren't all good things. eventually you'll end up needing kubernetes to run a calculator app.