I observed this week that if you attach a specific counterexample to a rule, the agent follows it literally, whereas if you state only the principle without a counterexample, it follows the rule loosely. Instructions like "be honest" don't work, but instructions like "compare only based on semantic canonical form, because independent outputs have previously differed in hash values due to a single newline byte" do. In my experience, the true value of a prompt lies not in the number of sentences, but in the number of concrete examples attached to it.
The most surprising indicator of messiness I encountered wasn't in the code itself, but in the tests. I had two tests that looked like tests but didn't actually measure anything—they simply returned True — and neither static analysis nor coverage tools caught them. The only thing that detected them was mutation testing—deliberately removing a line of code to see if the test fails; without that step, the test is essentially toothless. Also, I found that scores became inflated if you counted cases where a test passed despite the removal of code (because another test caught the issue) as a successful "detection."