Compactly formatted user messages are something an agent can ingest in a few minutes, even if they are thousands of lines long. And the quality of those messages is great: they don't track what the agent does well, only what changes and what breaks.
Having this top-down view helps a lot. Usually, within a session and deep into a task, the agent loses the global perspective and optimizes for local success. I find it weird there is no harness that treats user messages as high value signal (except my own, of course, I have it, https://github.com/horiacristescu/playbook-harness).
However, I'd bring in Brooks discussion on essential and accidental complexity. In other words, there being No Silver Bullet https://www.cs.unc.edu/techreports/86-020.pdf
The problem with specification and user discussion is they still have errors that code has. But unlike code, there are no tests to confirm correctness.
So now we have a definition of the system in a non-exact language with no ability to test to confirm it's correctness. The code holds the essential complexity and now we are adding accidental complexity on top to manage.
Again agree the specifications and user discussion provides context for the AI. However, a well written test suite provides similar context that can actually confirm correctness of the system.
However, saying all the above. Focus of ImpactGate ( https://impactgate.officefloor.net ) is about erosion of the code, not correctness.
The code becomes a reflection of that.
I'm interested in your experiences of capturing specifications and user discussion on whether this captures the end intentions? Or whether it keeps you focused on earlier dead end directions?
Besides intent, I also mine signs of "user friction," which I use as input for the agent to come up with new tests. What I complain about is one of the signals driving testing.
I'd be interested to see what happens:
- to token counts after a year of so of changes, as the specification list grows?
- how it goes with concurrent changes in teams?
Plus whether asking AI to add good commenting to the code could achieve the same thing?
Do you mean like the hyper focus an LLM puts on the task in front of it so you end up with drift (duplicated concepts/multiple ways of doing things, terminology drift (e.g., now we have "customer" and "client"). That sort of thing?
Yes, there are generally complex algorithms but they usually are not things developers write (imported from libraries).
What is usually going on in the god class/method is that things keep getting added to it. These things should be separated out. So the cohesiveness of the class/method erodes into doing too many things.
The idea of the Change Impact formula is to catch this early so you start refactoring to separate out into classes with single cohesive purposes.
The problem with AI is it handles complexity really well and will happily keep piling changes into god classes/methods reaching ridiculous CC levels (have see over 200). Previously developers would get annoyed and do the refactor. But with AI these days, changes are happening faster. So Change Impact is to try to monitor the cohesive erosion.
High cohesion means the functionality of a component are closely related and focused on performing a single well defined task. Basically single classes for single purposes.
Erosion of this is when classes start doing to many things, in the case of God classes.
The Change Impact formula looks at a way of detecting when the cohesion is eroding and flagging it on a change (as the pull/merge request itself should generally be single focus cohesive change)
While a lot of metrics make intuitive sense, we don’t have that much hard evidence to prove or disprove their value. Part of it is the whole “if a metric becomes a target, it ceases to be a good metric” thing. Adding the checks to a large existing project probably has negative value. But I think it’s worth doing for greenfield projects.
For humans, these should just be advisory. But for LLMs I’m happy enough to make it a blocking check.
I keep thinking of doing an experiment where I give the same LLM the same problem, and only change which metric is enforced. And then see if any of them have a noticeable effect on correctness/maintainability.
> any examples of how people get this into an actual report / CI test / benchmark / whatever ?
Yeah they have examples of adding it to CI, or local checks, generate html reports, etc in their docs.
I've been in the process of reviewing and validating a lot of tools like this (qlty, Sonarqube, fallow, etc) and the false positive rate is anywhere from 20% to 80% for a lot of our sniff tests (zizmor produces an overwhelming majority of false positives here for what feels like arbitrary and very context-dependent GHA requirements)
the last thing I want to do is to annoy the hell out of our devs by requiring checks like these to pass especially since it's only a small percentage of them who vibe code everything and then also vibe response to code reviews. I feel like that's the anti-pattern that we'd push people towards by requiring checks like these to pass
another avenue of exploration has been requiring test coverage but also good test quality metrics (eg are there negative tests? mutation testing? empty asserts?) something that seems quite easy to spin up into a skill and pair with a deterministic harness. trash-tests is a neat little project that incorporates some of this: https://github.com/frangelbarrera/trash-tests (disclosure: I am not the repo owner or even a contributor, just a quality nerd who loves underdogs lol)
all in all, it really does feel like we'll need a revamp of the SDLC with our current expected velocities
Though to the bigger point of your comment, yes SDLC are becoming faster. We can churn out code at a ridiculous rate. However, doesn't mean it's good code. And hence, there are studies showing things actually slowing down because reviews pile up.
I guess I look at the Change Impact formula (and https://impactgate.officefloor.net implementation of it) as threshold tool. Small changes that aren't contributing to god classes, just let through. When things start to smell, the files involved get marked for review.
Ideally this then can cut down on review time and allow overall increased velocity.
But yes relies on trusting the AI to do "simple" things
I was testing the additive pipeline style of OfficeFloor against the mutative handler style of Spring. I was looking to see what factors could be used to allow AI to make long on going changes (experiment is 60 changes to an end point, where all add functionality and every 4th change is mutative on existing rules). Then I watch how AI manages to make the 60 changes in each architecture.
I've done many runs and you are quite right about Goodhart effect in giving it the metric. Never knew Spring code could be written so badly.
I've tried runs with better prompting also and I'm starting to find the key factor is actually the architecture itself.
From my initial findings, it's seeming that additive pipeline architectures hold up much better against AI slop than our typically single method web handler architectures.