As far as I understand, inspecting the compiled binaries in any way would violate clean-room principles. At a minimum you'd need (A) one set of agents documenting the behavior of the executing binaries, e.g. manipulating the UI, and reviewing documentation and training manuals, and (B) a second set of agents whose only input is the output of the first set of agents. I'm not sure whether testing the real and AI generated agent side by side with equivalent inputs would constitute clean-room, although it might be legal.
I agree that a large part of the value of a product stems from "thousands of tiny design, engineering, and product decisions." But I don't agree with your conclusion. If the goal is to replicate the behavior of an existing piece of software then all you need to do is observe the behavior and replicate it. You don't need to engineer it in the same way, or arrive at it by making the same engineering and design decisions. Of-course in the end all you have is the replica, not the engineering and design heritage, product management, or the teams to carry the work forward. But that is the same problem open source "clones" of commercial products have often suffered from.