Long story: we have a big legacy desktop app. It uses a big legacy UI component (a grid control), which we had a license for in an old version. Fast forward 20 years, and to be able to move to a new runtime for our app, we need to update the component. Someone had bought the company making the component and now charges north of $1k per developer per year. So instead of doing this, we had just lived with the very old version.
We had long thought of writing our own control to replace the proprietary, but it was always going to be a man-year of work we thought. But I thought I'd give it a try with AI now. I told Opus: look at our uses of that control (tens of thousands of lines of code, it has over 100 instances across our User Interface). Write a new control that would compile with the exact same app syntax. First just make a dummy implementation that throws on every call. Then start implementing. Make a test suite that can run both with our new control and the proprietary control, and test everything, every function that can be called in its public interface and every state that can be inspected from the public API. Verify that everything behaves exactly the same, and lock it in with thousands of tests. Finally, check that the control _looks_ exactly the same as the proprietary one. Render to bitmaps, figure out the rendering logic from observation, such as arithmetic for padding, font sizes, and so on. Compare pixels until it's exactly the same.
Basically: it was a mammoth coding task, but it was so extremely well specified that an LLM could easily just do it. It's a clean-room implementation of something with no tests, but we had a test double that could provide 100% of the expected behavior. The description was extremely short. "Make a new thing that works like the old thing, and prove that it does". Opus 5.5 finished this in a number of hours. 500 source files, several thousand unit tests, and html reports with image diffs from the reimplementation and the original control. It did not use any disassembly or such "cheating". Only observation of the public API and the behavior. Do we need to deeply understand the implementation? Does the architecture matter? Not much in this case I'd argue. It was a black box to begin with and it remains a black box. If we notice a bug, we can always point it to the original proprietary control and say "there's a behavioral difference when doing X" and it will fix it, and lock it down with tests.
As a programmer it's kind of chilling. I had recreated for a few tens of dollars something that would cost $1000 per year to buy. Obviously it's not a complete implementation only the parts of the API we use. It likely still has some bugs. We don't get support, we get to maintain it ourselves. But the rate of reverse engineering this thing "black box" was frightening. It hasn't created anything novel. But we must realize that as programmers some times we have man-years of work that just isn't novel. And in the past, we didn't do this work at all.
I wonder if those who write and sell libraries like this will start having explicit no-reverse-engineering EULAs soon? Perhaps even explicitly mentioning AI/LLM use in analysis and reimplementation?_ Obviously the library we reimplemented was from 2005 so didn't mention AI... (It doesn't mention reverse-engineering either, luckily).