This looks very useful! What type of sandbox are you using? How does the mocking work?
OSX is the best, we use the in-built (seatbelt) sandbox via sandbox-exec.
For Windows, we use WSL containers when available.
By default, if a safe sandbox environment is not available, we inform the user that a repro is not possible in the current conditions.
We do a lot of AST parsing - for both code and build configuration languages. Even then, we still have to rely on the LLM to figure out a lot of the details.
Making this work reliably for non-frontier models and codebases that don't have existing test harnesses is where a lot of the design work goes in.