36 karma · joined May 13, 2025
---> Our pentest tool has a "secret" step called verification. We run a second agent to verify all the findings are "real". We have a built some pretty complex backend harness on top of our open-source mcp-xray to automate testing. If you are at the DEF CON this year, come to our demo labs and we can chat more.
Would be great to hear more about how you maintain coverage on changing AI/MCP combos without constant manual tuning. ---> It's very hard to be honest. We use agents everywhere but manual tuning is still needed.
Today, mcp-xray (https://github.com/traceforce/mcp-xray) can be used to dynamically (not just statically) test against your MCPs. It can detect security vulnerabilities such as code execution, SSRF, path traversal, authorization bypass, input injection, DoS etc. You can find us at demo labs during DEF CON this year. That said, we think behavioral influence is an important problem and it's an area we're interested in exploring next.
And I know Harold well at Bluerock. I totally agree that the market is crowded and will only get more crowded which is a good sign that the problem is real.
And yes I totally agree with you that differentiation is the key. I can't say that we have figured this out 100% but our approach is always community first, open-source first. I hope that is the right direction in the long run.
The challenge we hear from customers is that they don't know what AI apps, MCPs, or tools their employees are actually using. And new things just keep popping up everyday. Without that visibility, it's difficult to know where to apply controls.
I agree that DROP TABLE executes remotely. The key point is that the decision to invoke the tool is made by coding agents like Claude Code. Traceforce captures those tool calls at the application layer before they're executed.
The gap we focus on is application-level visibility inside AI apps—understanding which MCPs, skills, and tools are connected and what they're doing. A big part of our work has been building an MCP registry so we can accurately identify and classify MCPs, something traditional EDR telemetry doesn't provide.
That said, if your existing EDR already gives you that level of visibility and enforcement, I’d be interested in learning about which EDR you’re using and how it handles MCP and tool-level activity.