In my experience, spec drift is the main reason why none of these tools work. Maybe they work for one shot greenfield feature generation but in a large, multi developer long lived code bases, specs rot and end up being more pain than they are worth.
In my experience, spec drift is the main reason why none of these tools work. Maybe they work for one shot greenfield feature generation but in a large, multi developer long lived code bases, specs rot and end up being more pain than they are worth.
1. Lack of closed loop between the "spec" and working code (your spec rot point). The Rational Unified Process(RUP) was a grossly inefficient, heavily manual undertaking. Mapping between artefacts - e.g. "Platform Independent Models" and "Platform Specific Models" was a manual, largely heuristic based approach. As a consequence the models were not generally kept up to date as the project evolved.
2. User experience mismatch. Developers were asked to create diagrams instead of writing code. Tool usability was poor ("write code with a mouse") and the artefacts didn't fit well with necessary tools like diffing and source code control (try diffing an xml file textually).
Coding agents have some potential for alleviating (1) in that they can read the result code and, at least to some extent, ensure spec and code are in sync.
(2) is more open. Some users - those proportionally more interested in solving the problem than designing/writing code - are more comfortable with natural-language-based specs and exploration. Those more experienced/comfortable with code will likely see those specs more akin to UML diagrams: a distraction from the real thing.
I believe writing specs is different with AI for a few reasons: (1) natural language is expressive enough and the team collaborates at this level already, (2) LLMs can fill in the gaps, point out inconsistencies, and reliably map natlang to code, and (3) LLMs can read and refine specs at superhuman speeds, which makes spec maintenance economically viable for the first time ever outside of high stakes applications.
For this to work over the long run, specs must take a certain form. IMO: they must focus on original intent and what must be true after implementation (assertions) rather than implementation details. I also don't think one needs to specify anything an LLM can easily infer, so specs should be kept lean.
For spec drift, my team uses a sandboxed agent that checks for drift daily, triages, and surfaces issues. Beyond fixing specs, this has revealed a lot of product level miscommunications and helps us get ahead of them.