- LLMs struggle, or at least burn oodles of tokens, on big complex codebases. Managing tech debt helps the agents as well IMO. - If you ask an LLM to review a non-trivial pull request, it nearly always seems to find at least 10 issues. Many of which are not worth pursuing. Often, they would introduce a lot of extra complexity to harden the code against scenarios we don't care about or can never happen. So there's a lot of room for human judgement there. Typical scenario - the LLM reviewer finds 10 issues and I judge that 2-3 are worth pursuing - As you said, often the architectural approach sucks even with frontier LLMs. Usually from a lack of problem space / usage scenarios more than a lack of technical chops - LLMs tend to err on the side of overengineering the shit out of everything
I do not see how LLMs will be able to manage huge, hairy polygot codebases without humans-in-the-loop in the near future.
Given the pace of progress, I'd be a fool to bet against LLMs in the medium term future. But I think it's far from a given. LLMs themselves, and large codebases, are essentially many-to-many problems approaching something like O(n^n).