Thanks for trying it! So far, the the testing and validation stage has been left to the developer. While this could change in the future, my experience has been the models aren't quite good enough yet to make an auto-test/auto-fix loop like you've described the default behavior. You end up with a lot of tire-spinning where you're burning a lot of tokens to fix issues that a human can resolve trivially.
I think it's better for now to use LLMs to generate the bulk of a task, then have the developer clean up and integrate rather than trying to get the LLM to do 100%.
That said, you can accomplish a workflow like this with Plandex already by piping output into context. It would look something like:
plandex new
plandex load relevant_context.ts some_more_context.ts
plandex tell 'some kind of complex task'
# ...Plandex does its thing, but doesn't get it 100% right
npm test | plandex load
plandex tell 'please fix the problems causing the failed tests'
As the models improve, I'm definitely interested in baking this in to make it more automated.