I am surprised we haven't had an actual Y2K crash with these AI codes. Like how do you review a 1000 lines of Claude generated PR?
I remember having to do such a code review before an AI in a highly complex component, and it would take a full day of work to do it. In this day and age, most of the people i know take like half an hour and are mostly scanning for obvious mistakes, where the bigger problem are those sneaky non obvious ones.
What could work - llm creating a very good test suite, for their own code changes and overall app (as much as feasible), and those tests need a hardcore review. Then actual code review doesn't have to be that deep. But if everybody is shipping like there is no tomorrow, edge cases will start biting hard and often.
Pro tip: tell it the code came from Claude. That will make it put its war face on.