You can get the LLM to run a script which checks for all of these and also enforce them by running the same script as a pre-commit hook. Setting this up religiously in every code base I work on has been what's given me the most mileage with agentic coding.
I wrote down a more detailed post of the various linters I use here:
https://www.balajeerc.info/Use-Deterministic-Guardrails-for-...
I have legacy endpoints that are no longer used in practice, there for historical reasons, intertwined with existing code etc. They might be marked obsolete, services implementing it are not - agent greps those, builds off of them - produces half legacy garbage.
Linters only handle trivial cases most of us already solved.
// LEGACY CODE, per docs/legacy_rules.md §14, §19Most of the time. Except for when it forgets to do it.
I think it’s funny that the solution is to use something that is not LLM driven to enforce it.
Also - pre commit hooks aren’t enforced, people will not set them up. You have to run this stuff in CI (which is incredibly annoying given that machines are writing the code in the first place)
Putting structural code checks in a precommit hook is arguably better than pulling it into the harness, as it will enforce those constraints no matter whether an agent or human is making the commit.
> what if your harness hook malfunctions?
That’s a bug in the harness and should be enforced
Unless the agent, or the human don't enable the precommit hooks in the first place.
Adding an auto runner of the unit tests is just... Boring?
The reason is quite obvious if you have dealt with such a huge code basis in production with thousands of developers contributing for decades to it coming from different vendors and countries.
Code has a meaning attached to it. And paradoxically being able to cleanly cut out such dead code raises my suspicion. There is a reason why such code exists in there often times and since almost always stakeholders give a damn about documentation, and developers traditionally have a hard time writing even JavaDocs, JSdocs, whatever and not to mention maintaining them.
In earlier times CPU time was precious and comments were deliberately left off due to space and processing considerations.
So why is this all important?
Because until you cannot find the one guy who uses this code for a good reason, I would never kill it. Good reasons in these cases are almost always so called application owner, an app admin and hosts, who serves according to ITIL specs as deployment and production person.
Dead code can be actually quite lively under the right circumstances. And since sometimes people have to be very creative to serve regulation requirements and compliance, sometimes release pressure or missing tools can make such code an important script or deployment tool or fix for a reboot or whatever.
Believe me, dead code isn't. What you can do is, watch it at least over a period of two years.
Here is why.
Most processes have yearly deadlines. Many fall on the 1.1. of each year, while others somewhere at the end of the year. Some processes need to be served once a year due to compliance to laws.
And why two years then? As I said, human beings. Maybe that one time a guy had an exception running for it or there is a maintainer, who uses a different method - his own dead code so to say - because these scripts are rarely shared and maintainer's best kept secrets. The lesser a company knows, the more important these last line folks feel and they are blackboxes when something is or isn't working. (I hated this, this was not way of working and I changed it. It is not their company and a keeper is a Red Flag for me.)
But rarely the same person will serve the process two times in a row. Vacation times vary as well as positional changes.
Hence the two years period and even then, there are smarter, more easier ways than to use linters.
Hint: it is the frontend first paradigm.
How do I know? Because I invented it. Proof: Huge international Bank. dbCORE.
You’ve never had an agent completely lose the plot and forget/confuse its instructions due to the context filling up?
I guess it depends on what you mean with "better" but almost all the agent-built projects I do with zero regards to code quality, design and architecture ends up with every single agent needing 10+ minutes to do even the easy changes, while the ones where I focused on those things together with the agent, large changes can take 10+ minutes but everything else is solved faster.
I don't have empirical evidence of this yet, I guess I should put together some sort of test to confirm/disconfirm this.
The most useful discussion would be if we all read the paper and critique its methodology or results.
though people who complain that llms aren't that great strike me as the type to have messy code bases