My org has built internal tooling that approximates this. It's incredibly valuable from a manual test perspective though we haven't managed to get the agent part working well, app startup times (10+ min) make iterating hard.
Do you have customers who have faced/solved this problem? If so, how did they do it -- it seems like a killer on the approach?