how do you evaluate your tool, and have you published your evaluation along with the metrics?
Our key metrics include the time and cost per agentic loop, as well as the false positive rate for a full end-to-end test. If you have any specific benchmarks or evaluation metrics you'd suggest, we'd be happy to hear them!
I’m not aware of any evals or shared metrics. But measuring a testing agents performance seems pretty important.
What is your tool’s FPR on your golden suite?