Don't get me wrong, this is impressive and capabilities do still seem to be on a generally upward trend with nothing "hitting a wall" despite all the parroting of that phrase, but there's a huge gap between the first time a machine manages something impressive enough to document, and that machine becoming so good at that task that humans need not apply.
"Is this the end of human code review?" != "the end of human code review"
Now since the problem in this study is more complex than 90 to 99.5% of what a normal ticket is, the question makes sense. The agent does not have knowledge of the project, at the begining of each session, and manages to handle something more complicated than virtually anything a developper has to do. The question holds IMO.
The fact is, the agent did something amazingly difficult, that zero human on earth could do. Read the spec if you believe someone could do it. It did it clean, everything is public, and it works perfectly. Zero code review.
https://aisovereignlabs.ai/docs/case-study/liveSession/logs/...
99.9% of the IT dev tickets are easier than this, excluding off topics like discrete math, computer vision algorithms, etc, which are irrelevant here.
Do you think there's anything you've done, with human code review, in the past 30 years, matching the difficulty of this ?