The output of an LLM is a distribution, and yes, if you’re just taking the first answer, that’s problematic.
However, it is a distribution, and than means the majority of solutions are not weird edge cases, they’re valid solutions.
Your job as a user is to generate multiple solutions and then review them and pick the one you like the most, and maybe modify it to work correctly if it has weird edge cases.
How do you do that?
Well, you can start by following a structured process where you define success criteria as a validator (eg. Tests, compiler, parser, linters) and fitness criteria as a scorer (code metrics like complexity, runtime, memory use, etc)… then:
1) define goal
2) generate multiple solution candidates
3) filter candidates by validator (does it compile? Pass tests? Etc)
4) score the solutions (is it pure? Is it efficient? Etc)
5) pick the best solution
6) manually review and tweak the solution
This structured and disciplined approach to software engineering works. Many of the steps (eg. 3, 4, 5) can be automated.
It generates meaningful quality code results.
You can use it with or without AI…
You don’t have to follow this approach, but my point is that you can; there is nothing fundamentally intractable able using a language model to generate code.
The problem that you’re critiquing is the trivial and naive approach of just hitting “generate” and blindly copying that into your code base.
…that’s stupid and dangerous, but it’s also a straw man.
Seriously; people writing code with these models aren’t doing that; when you read blogs and posts from people, eg. Building seriously using copilot you’ll see this pattern emerge repeatedly:
Generate multiple solutions. Tweak your prompt. Ask for small pure dependency free code blocks. Review the and test output.
It’s not a dystopian AI future, it’s just another tool.