Genuinely: how? Share some insights. So far I have seen anything exceeding about 2k lines of text in the context diverge. Meaning, throwing more LLM at it rapidly baloons the size at the expense of internal coherence. Things become stale, hallucinated, duplicated and outright faked. Only thing that works is constant manual intervention and pruning of the generated slop.
Not only these things are unable to find "best solution", they are unable to find any solution. Instead (especially for Opus) they seem optimized to convince user that the task is accomplished.