I've looked into developing "digital twins" of APIs that I rely on, similar to the StrongDM pattern, but the models talk me out of it every time. The rationale is that there could be subtle differences or drift between the behavior of the real service and the mocked API. So you may think you're on solid ground, but once you switch to prod, you could see show-stopping differences.
I think if I was doing evals on the way an agent interacts with a service (e.g. Slack) the digital twin could make sense, but for rapid iteration or testing, it might not be what you want.