The closed-decision part is useful. Re-solving the same decision in every run feels like one of those costs that isn't obvious until you have a lot of agent traffic.
The core vs plugins split gets more important once the agents start changing the system too. Otherwise the harness itself becomes another thing the agents have to understand before they can do any useful work.
I can see the value of having the agent and the human looking at the same design. The thing I'd worry about is the diagram becoming another thing that slowly diverges from the code. Would be interesting to see how you handle that once the implementation starts moving.